Biasing the interpretation of spoken language received in a vehicle environment.

By biasing the interpretation of oral utterances in vehicles using vehicle-specific sensor data, the system addresses misinterpretation issues, enhancing accuracy and reducing resource consumption.

JP7848345B2Active Publication Date: 2026-04-20GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2022-06-30
Publication Date
2026-04-20

AI Technical Summary

Technical Problem

Automated assistants in vehicle environments often misinterpret spoken utterances related to vehicle operations due to a lack of context awareness, leading to incorrect actions and increased user input and resource consumption.

Method used

Bias the interpretation of oral utterances in a vehicle environment by utilizing vehicle-specific sensor data to restrict or prioritize searches within vehicle-specific user manual corpus data, using criteria such as temporal relationships, user association with the vehicle, and explicit instructions.

Benefits of technology

Reduces computational resources and improves accuracy by limiting search space to vehicle-specific data, ensuring correct responses to user queries related to vehicle operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007848345000001
    Figure 0007848345000001
  • Figure 0007848345000002
    Figure 0007848345000002
  • Figure 0007848345000003
    Figure 0007848345000003
Patent Text Reader

Abstract

Implementations described herein relate to various techniques for biasing interpretations of verbal utterances received in a vehicular environment. For example, implementations may receive a verbal utterance including a query from a user of a vehicle and obtain corresponding vehicle sensor data instances generated by a vehicle sensor(s) of the vehicle. Some implementations may determine to perform a search only on the first corpus data and not on the second corpus data to obtain a given response to the query based on various criteria, which may include at least the query, the corresponding vehicle sensor data instances, corresponding timestamps associated with the corresponding vehicle sensor data instances, and / or corresponding durations the user is associated with the vehicle. Additional or alternative implementations may perform a search on both the first corpus data and the second corpus data to obtain a given response based on the criteria.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans may engage in human-computer interaction using interactive software applications referred to herein as “automated assistants” (also known as “chatbots,” “conversational personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversational agents,” etc.). For example, a human (sometimes referred to as a “user” when interacting with an automated assistant) may provide oral natural language input (i.e., utterances) directed to the automated assistant, which may in some cases be converted to text and then processed, and / or provided as textual (e.g., typed) natural language input directed to the automated assistant. These oral utterances and / or typed inputs often contain assistant commands directed to the automated assistant. The automated assistant typically responds to these assistant commands by providing responsive user interface output(s) (e.g., audible and / or visual user interface output), controlling smart devices(s), and / or performing other actions(s).

[0002] These automated assistants typically rely on a pipeline of components to interpret and respond to these spoken utterances and / or typed inputs. For example, an automatic speech recognition (ASR) engine can process audio data corresponding to the user's spoken utterances and generate ASR output, such as a transcript of the utterances (i.e., a sequence of words and / or other tokens). Furthermore, a natural language understanding (NLU) engine can process the ASR output (or typed input) to generate NLU output, which may include the user's intent in providing the spoken utterances and, optionally, slot values ​​for parameters related to that intent. Additionally, a fulfillment engine can be used to process the NLU output and generate fulfillment output, such as a structured request to retrieve content that responds to the spoken utterances.

[0003] In some cases, these automated assistants may be used in specific environments. For example, a given automated assistant may be associated with the in-vehicle computing device of a user's vehicle. In this example, the given automated assistant can utilize the pipeline of the components described above when interpreting and responding to these spoken utterances and / or typed inputs. However, in some of these cases, the given automated assistant does not take into account that it will be used in a vehicle environment when interpreting and responding to these spoken utterances and / or typed inputs. As a result, some spoken utterances associated with vehicle operations may be misinterpreted, leading to incorrect actions or the provision of incorrect response content. Furthermore, the user may provide additional spoken utterances and / or additional typed inputs to ensure that the correct actions and / or response content are provided, thereby increasing the amount of user input and wasting computing resources. [Overview of the project]

[0004] The embodiments described herein relate to various techniques for biasing the interpretation of oral utterances received in a vehicle environment. For example, an embodiment can receive an oral utterance from a user located in the user's vehicle, including a query, and obtain a corresponding vehicle sensor data instance generated by the vehicle's vehicle sensor(s). In some embodiments, it can be decided, based on at least the query and the corresponding vehicle sensor data instance, to perform a search only on a first corpus data and not on a second corpus data to obtain a given response to the query. In these embodiments, the search can be performed on the first corpus data to identify one or more candidate responses, and a given candidate response from among the one or more candidate responses can be provided to the user in response to the query. In other embodiments, the search can be performed on both the first and second corpus data to identify one or more candidate responses, and a given candidate response from among the one or more candidate responses can be provided to the user in response to the query. In these embodiments, the selection of a given candidate response can be biased towards any candidate response obtained from the search on the first corpus data based on the corresponding vehicle sensor data instance. In particular, the first corpus data in these embodiments can correspond to user manual corpus data specific to vehicles and provided by vehicle original equipment manufacturers (OEMs). Furthermore, the second corpus data in these embodiments can correspond to any corpus data that is not specific to vehicles, such as web-based corpus data.

[0005] For example, suppose we receive the verbal utterance, "Which tire has low pressure?" from the vehicle user while they are inside the vehicle. Furthermore, suppose we obtain a tire pressure sensor data instance from the vehicle's tire pressure sensor, indicating that the passenger-side tire of the vehicle has low air pressure. In some examples, the tire pressure sensor data instance can cause the vehicle's dashboard light to illuminate and / or provide an alert to be presented to the user on the display of an in-vehicle computing device to make the user aware that the passenger-side tire has low air pressure. Thus, the interpretation of the verbal utterance provided by the user can be biased towards vehicle-specific user manual corpus data. As a result, a given candidate response to the query, "The tire pressure is 28 psi, but according to the user manual it should be 32 psi," can be selected to provide the user with an audible and / or visual presentation in response to the query. In particular, in this example, the given candidate response may be selected in preference to one or more other candidate responses using various bias criteria. Without these various bias criteria, not only could vehicle-specific user manual corpus data be searched, but one or more other non-vehicle-specific corpus data could also be searched.

[0006] In some embodiments, this bias determination can be based on determining whether an oral utterance containing a query was received within a threshold duration for which a corresponding vehicle sensor data instance is generated, and / or as indicated by the corresponding timestamp associated with the corresponding vehicle sensor data instance. The threshold duration may include, for example, a static duration (e.g., 30 seconds, 2 minutes, etc.) or a dynamic duration based on other factors. For example, suppose, as described above, the tire pressure sensor data instance illuminates a light on the vehicle's dashboard and / or provides an alert to present to the user. Furthermore, suppose the oral utterance is received within 15 seconds after the dashboard light illuminates and / or the alert is displayed. In this example, it can be determined that the oral utterance contains a query associated with the dashboard light illuminating and / or the alert being displayed. This, therefore, makes it possible to utilize the temporal relationship between the provided oral utterance and the corresponding vehicle sensor data instance in the bias interpretation of the oral utterance.

[0007] In an additional or alternative embodiment, this bias determination can be made based on determining whether an oral utterance containing a query is related to the corresponding vehicle sensor data instance. For example, again, the corresponding vehicle sensor data instance corresponds to a tire pressure data instance indicating that the air in the passenger-side tire is low, the dashboard light of the vehicle lights up as described above, and / or an alert provided for presentation to the user is provided, and it is assumed that an oral utterance of "Which one has low air pressure?" is received. The audio data capturing the oral utterance containing the query can be processed using an automatic speech recognition (ASR) model to generate one or more recognized terms corresponding to the query. Thereby, considering the dashboard light and / or the alert, the recognized term "low air pressure" can be determined (e.g., using software matching, semantic word matching, and / or other techniques) to be related to the passenger-side tire with low air. Thus, this enables the use of the linguistic relationship between the provided oral utterance and the corresponding vehicle sensor data instance, additionally or alternatively, for the biased interpretation of the oral utterance.

[0008] In additional or alternative embodiments, this bias determination can be based on determining whether the user who provided the oral utterance was associated with the vehicle for a threshold duration. The threshold duration may represent, for example, the amount of time the user spent in the vehicle, the distance the user drove the vehicle, the number of times the user started the vehicle, the number of times a given corresponding vehicle sensor data instance generated by one or more sensors of the vehicle was captured, and / or other factors. For example, when processing audio data that captures oral utterances containing queries, speaker identification of the audio data can be performed to determine whether the user who provided the oral utterance is a known user. Speaker identification can be performed using any known technique (e.g., text-dependent speaker identification, text-independent speaker identification, etc.) and any known speaker identification model. Additional or alternative techniques can be used to determine whether the user who provided the oral utterance is a known user, such as facial recognition based on processing visual data generated by a visual sensor(s) of the computing device or an additional computing device, fingerprint identification based on processing fingerprint data generated by a fingerprint sensor(s) of the computing device or an additional computing device, and / or any other technique.

[0009] Furthermore, the identification information can be compared with the identification information of known users of the vehicle (e.g., accounts associated with the vehicle or the automatic assistant) to determine whether the user is a known user. In these embodiments, if the user is a known user but not associated with the vehicle for a threshold duration, the interpretation of the spoken utterance can be biased towards the user manual corpus data. Additionally or alternatively, if the user is not a known user, the interpretation of the spoken utterance can be biased towards the user manual corpus data. However, in these embodiments, if the user is a known user and is associated with the vehicle for the threshold duration, the user is likely to already be well aware of the reason for the vehicle dashboard lighting up and / or the reason for the alert being displayed (e.g., the tire pressure on the passenger side is low), so the interpretation of the spoken utterance may not need to be biased towards the user manual corpus data. Thus, this enables the duration for which the user is associated with the vehicle to be used, additionally or alternatively, when biasing the interpretation of the spoken utterance.

[0010] In additional or alternative embodiments, this bias determination can be made based on determining whether the spoken utterance includes an explicit instruction to search only the first corpus data. For example, again, assume that the corresponding vehicle sensor data instance corresponds to a tire pressure data instance indicating that there is less air in the tire on the passenger side, and assume that the vehicle dashboard light is on and / or an alert is displayed as described above. However, assume that a spoken utterance is received such as "Check the user manual and confirm the one with low pressure" (rather than simply "Which one has low pressure"). In this example, the query included in the spoken utterance also includes an explicit instruction to search only the first corpus data as indicated by "user manual". Thus, this enables the user to provide, additionally or alternatively, an explicit instruction that can be used when biasing the interpretation of the spoken utterance.

[0011] While the above technologies are described as being implemented in a vehicle's on-board computing device, it should be understood that this is for illustrative purposes only and not intended as an limitation. For example, the technologies described herein can be additionally or alternatively implemented in a vehicle user's mobile computing device. Furthermore, the technologies described herein can be additionally or alternatively implemented by a remote computing device (e.g., a remote server or a cluster of remote servers). However, in various embodiments, on-device processing by the on-board computing device and / or mobile computing device may be preferred to reduce latency when responding to queries contained in spoken utterances.

[0012] Furthermore, while the above techniques describe the use of corresponding vehicle sensor data instances in determining which of one or more corpus data to search and / or performing searches on one or more identified corpus data, it should be understood that this is also illustrative and not intended to be limiting. For example, one or more of these searches may be performed without using any corresponding vehicle sensor data instances. For example, some vehicles do not have an in-vehicle computing device capable of implementing the automated assistant described herein and / or a Controller Area Network (CAN) bus for obtaining the corresponding vehicle sensor data instances. Nevertheless, in these examples, an automated assistant implemented on the user's mobile computing device may use the same or similar techniques as those described herein to bias the interpretation of spoken utterances by limiting the search space for queries associated with the user's vehicle.

[0013] By using the techniques described herein, various technical advantages can be obtained. As a non-limiting example, the techniques described herein enable the system to efficiently restrict the search space for identifying candidate responses to a query based on queries for vehicle sensor data generated by vehicle sensors and / or corresponding vehicle sensor data instances. For example, the techniques described herein can bias the interpretation of spoken utterances by performing a search on a given corpus data (e.g., vehicle-specific user manual corpus data) based on a query to identify a given candidate response to the query, and by avoiding the use of other corpus data based on various contextual signals and / or data. Alternatively, for example, the techniques described herein can bias the interpretation of spoken utterances by performing searches on multiple corpus data (e.g., vehicle-specific user manual corpus data and at least one additional corpus data that is not vehicle-specific), but by biasing the selection of a given candidate response to a query based on what is obtained from the given corpus data. As a result, the consumption of computational resources when processing spoken utterances and identifying a given candidate response can be reduced.

[0014] The above description is provided as a summary of only some of the embodiments disclosed herein. These embodiments and other embodiments will be described in further detail herein. [Brief explanation of the drawing]

[0015] [Figure 1] This document illustrates various aspects of the present disclosure and provides block diagrams of exemplary hardware and software environments that can implement embodiments disclosed herein. [Figure 2] This shows an exemplary process flow that biases the interpretation of oral utterances (or multiple utterances) received in various vehicle environments, as shown in Figure 1, according to various embodiments. [Figure 3]A flowchart is provided illustrating exemplary methods for biasing the interpretation of oral utterances received in a vehicle environment, using various embodiments. [Figure 4] A flowchart is provided illustrating another exemplary method of biasing the interpretation of oral utterances received in a vehicle environment through various embodiments. [Figure 5] A flowchart is shown illustrating yet another exemplary method of biasing the interpretation of oral utterances received in a vehicle environment, through various embodiments. [Figure 6A] This paper presents various non-limiting examples of computing devices that demonstrate various user interactions that bias the speech processing of oral utterances in a vehicle environment, through various embodiments. [Figure 6B] This paper presents various non-limiting examples of computing devices that demonstrate various user interactions that bias the speech processing of oral utterances in a vehicle environment, through various embodiments. [Figure 7] This document illustrates exemplary architectures of computing devices in various embodiments. [Modes for carrying out the invention]

[0016] Referring now to Figure 1, an environment in which one or more selected embodiments of this disclosure may be implemented is shown. An exemplary environment is a group of computing devices 110 1-N This includes an automatic assistant application bias system 120, a vehicle 100A, one or more original equipment manufacturer (OEM) applications 181, one or more first-party applications 182, and one or more third-party applications 183. These components 110 1-N,120, 181, 182, and 183 may communicate via one or more networks, for example, usually indicated by 195. One or more networks may include wired or wireless networks, such as local area networks (LANs) including Wi-Fi, Bluetooth, near-field communication, and / or other LANs, wide area networks (WANs) including the Internet, and / or other networks, in order to facilitate communication between the components shown in Figure 1.

[0017] In various embodiments, the user accesses the computing device 110 1-N One or more of these can be operated to interact with the other components shown in Figure 1. Computing device 110 1-N For example, desktop computing devices, laptop computing devices, tablet computing devices, mobile phone computing devices, and in-vehicle computing devices for vehicles (e.g., 110A). N This may include in-vehicle communication systems, in-vehicle entertainment systems, and / or in-vehicle navigation systems, or wearable devices including computing devices such as head-mounted displays ("HMDs") and "smart" watches that provide an immersive computing experience of augmented reality ("AR") or virtual reality ("VR"). Additional and / or alternative computing devices may be provided.

[0018] Computing device 110 1-NEach of the and bias systems 120 may include one or more memories for storing data and software applications (e.g., one or more OEM applications 181, one or more first-party applications 182, and / or one or more third-party applications 183), one or more processors for accessing the data and running the software applications, and other components for facilitating communication via one or more of the networks 195. Computing device 110 1-N The operations performed by one or more of the bias systems 120 may be distributed across multiple computer systems. For example, the bias system 120 may be implemented as a computer program that runs on only one or more computers in one or more locations that are connected to each other in a communicative manner via one or more of the networks 195, or it may run in a distributed manner across them.

[0019] Component 110 1-NOne or more of 120, 181, 182, and 183 may include various different components that can be used to bias the interpretation of oral utterances received in a vehicle environment, for example, as described herein. For example, computing device 1101 may include a user interface engine 1111 for detecting and processing user input (e.g., oral utterances, typed inputs, and / or touch inputs) directed to computing device 1101. As another example, computing device 1101 may include one or more sensors 1121 for generating corresponding sensor data. One or more sensors may include, for example, a GPS sensor for generating Global Positioning System ("GPS") data, a visual component for generating visual data within the field of view of a visual component, a microphone for generating audio data based on oral utterances captured in the environment of computing device 1101, and / or other sensors for generating corresponding sensor data.

[0020] As yet another example, the computing device 1101 can interpret various user inputs received by the computing device 1101 by operating an input processing engine 1131 (which may be, for example, standalone or as part of another application, such as an automated assistant application). For example, the input processing engine 1131 may capture spoken utterances and process audio data generated by the microphone(s) of the client device 1101 to generate an ASR output using an automatic speech recognition (ASR) model(s) (e.g., a recurrent neural network (RNN) model, a transformer model, and / or any other ML model capable of performing ASR). Furthermore, the input processing engine 1131 may process an ASR output (or type input) to generate an NLU output using a natural language understanding (NLU) model(s) (e.g., a long short-term memory (LSTM), a gated regressive unit (GRU), and / or any other type of RNN or other ML model capable of performing NLU), and / or grammar-based NLU rules(s). Furthermore, the input processing engine 1131 may process the NLU output to obtain one or more candidate responses in response to user input, such as actions that the automated assistant will perform based on user input, or content items that will be provided to present to the user based on user input, using fulfillment models and / or fulfillment rules. In embodiments where text content is audibly rendered in response to oral utterances or typed input, the user interface engine 1111 may process the text content to generate synthesized speech audio data, including computer-generated synthesized speech that captures the content, using text-to-speech models. The synthesized speech audio data may be audibly rendered for presentation to the user via the speaker(s) of the computing device 1101.In embodiments where visual content is visually rendered in response to oral utterances or typed input, the user interface engine 1111 can visually render the visual content for presentation to the user via the display 1101.

[0021] In various embodiments, the ASR output may include, for example, one or more speech hypotheses (e.g., terminology hypotheses and / or transcription hypotheses) predicted to correspond to the user's speech activity and / or oral utterances captured in the audio data, one or more corresponding predicted values ​​(e.g., probability, log-likelihood, and / or other values) for each of the one or more speech hypotheses, a plurality of phonemes predicted to correspond to the user's speech activity and / or oral utterances captured in the audio data, and / or other ASR outputs. In some versions of these embodiments, the input processing engine 1131 may select one or more of the speech hypotheses as recognized text corresponding to the oral utterances (e.g., based on the corresponding predicted values).

[0022] In various embodiments, the NLU output may include annotated recognized text, for example, one or more annotations of the recognized text for one or more (e.g., all) of the terms in the recognized text. For example, the input processing engine 1131 may use a part-of-speech tagger (not shown) configured to annotate terms with their grammatical roles. Additionally or alternatively, the input processing engine 1131 may use an entity tagger (not shown) configured to annotate entity references within one or more segments of the recognized text. Entity references may include, for example, references to people (e.g., literary characters, celebrities, public figures, etc.), organizations, places (real and fictional), etc. In some embodiments, data about entities may be stored in one or more databases, such as a knowledge graph (not shown). In some embodiments, the knowledge graph may include nodes representing known entities (and, optionally, entity attributes), as well as edges connecting the nodes to represent relationships between entities. Entity taggers can annotate references to entities at a high level of granularity (for example, to enable the identification of all references to an entity class such as a person) and / or at a low level of granularity (for example, to enable the identification of all references to a specific entity such as a particular person). Entity taggers may rely on the content of natural language input to resolve a particular entity and / or may optionally communicate with a knowledge graph or other entity database to resolve a particular entity. As described herein, NLU output can be used to determine whether an oral utterance contains one or more prominent terms corresponding to user manual corpus data.

[0023] Additionally or alternatively, the input processing engine 1131 can use a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more context queues. As a non-limiting example, the coreference resolver can be used to resolve the term "that" in a natural language input of "What is that light?" to a specific light or indicator associated with the operation of the vehicle 100A being generated, based on corresponding sensor data instances generated by one or more vehicle sensors (s) leading to that specific light or indicator associated with the operation of the vehicle 100A being generated. In some embodiments, one or more components utilized by the input processing engine 1131 can depend on annotations from one or more other components utilized by the input processing engine 1131. For example, in some embodiments, the entity tagger can depend on annotations from the coreference resolver when annotating all mentions of a specific entity. Also, for example, in some embodiments, the coreference resolver can depend on annotations from the entity tagger when clustering references to the same entity.

[0024] As yet another example, the computing device 1101 can operate a bias system client 1141 (which can be part of another application, such as stand-alone or part of an auto-assistant application) to interact with the bias system 120. Further, an additional computing device 110 N can be in the form of an in-vehicle computing device of the vehicle 100A. Although not shown, the additional computing device 110 N can include components the same as or similar to those of the computing device 1101. For example, the additional computing device 110 NThis may include each instance of a user interface engine for detecting and processing user input, one or more sensors for generating corresponding vehicle sensor data instances for vehicle sensor data, an input processing engine, and / or a bias system client for interacting with the bias system 120. In this example, one or more sensors are a tire pressure sensor for generating tire pressure data for the tires of vehicle 100A, an airflow sensor for generating airflow data for the air conditioning system of vehicle 100A, a vehicle speed sensor for generating vehicle speed data for vehicle 100A, an energy sensor for generating energy source data for the energy source of vehicle 100A, a transmission sensor for generating transmission data for the transmission of vehicle 100A, and / or the vehicle 100A and / or the vehicle-mounted computing device 110 of vehicle 100A. N It may include vehicle sensors, such as any other sensors integrated into it. Furthermore, Figure 1 shows computing device 1101 and in-vehicle computing device 110 N Only the following are shown, but please understand that this is for illustrative purposes only, and additional or alternative computing devices may be provided.

[0025] In various embodiments, the bias system 120 may include an interface engine 121, an input processing engine 122, a request processing engine 123, a user context engine 124, a vehicle context engine 125, a bias engine 126, a search engine 127, and a response engine 128, as shown in Figure 1. In some embodiments, one or more of the engines 121-128 of the bias system 120 may be omitted. In some embodiments, one or more of all or some of the engines 121-128 of the bias system 120 may be combined. In some embodiments, one or more of the engines 121-128 of the bias system 120 may be computed to a computing device 110 1-NIt may be implemented by a component that runs partially or completely remotely from one or more of the following. In some embodiments, one or more of the engines 121-128 of the bias system 120, or any operating part thereof, is the computing device 110 1-N It may be implemented in a component that runs partially or entirely locally by one or more of these components.

[0026] Referring to Figure 2, an exemplary process flow is shown that biases the interpretation of oral utterances(s) received in various vehicle environments as shown in Figure 1. Interface engine 121 (not shown in Figure 2) can facilitate data exchange with one or more of the engines 121-128 shown in Figures 1 and 2. For example, oral utterances can be transmitted from a user located inside vehicle 100A to a computing device (e.g., computing device 1101, in-vehicle computing device 1100). N Assume that it is received via (and / or any other computing device). The spoken utterance can be captured as audio data 201A and input processing engine 113 1-N / 122 can process audio data to generate processed input data 202. For example, audio data can be processed using an automatic speech recognition (ASR) model(s), a natural language understanding (NLU) model(s), a speaker identification model(s), and / or other models to generate processed input 202. Processed input data 202 may include, for example, ASR data, NLU data, and / or any other data resulting from the processing of audio data 202. The request processing engine 122 can receive the processed input data 202 and, based on analyzing the ASR data, NLU data, and / or other data contained in the processed input data, determine that the oral utterance contains a query. Furthermore, the request processing engine 122 can generate request data 222 that can be used when performing one or more searches on one or more corpus data as described herein. Furthermore, the request processing can provide the request data to the search engine 127.

[0027] In various embodiments, the input processing engine 113 1-NWhile / 122 and other components of the bias system 120 process the audio data 201, the bias engine 126 can process various signals and data acquired by the bias system 120 to generate bias data 226, and determine whether to restrict the query's search space to one or more specific corpus data. These various signals and data may include, for example, user context signals 224 acquired via the user context engine 124, vehicle context signals 225 acquired via the vehicle context engine 125, OEM (original equipment manufacturer) data 281 acquired via one or more OEM applications 181, first-party data 282 acquired via one or more first-party applications 182, third-party data 283 acquired via one or more third-party applications 183, and / or other various signals and data that may be used when determining whether to restrict the query's search space to one or more specific corpus data.

[0028] For example, user context signals 224 acquired via the user context engine 124 can characterize the user state of vehicle 100A. The user state of vehicle 100A may include, for example, whether the user is driving vehicle 100A, whether the user is a passenger in vehicle 100A, or the level of user involvement while vehicle 100A is in motion. The user state of vehicle 100A may additionally or alternatively include, for example, the user's starting point, the user's destination, the route the user is expected to take from the starting point to the destination, and / or any other signals that characterize the user state. Thus, user context signals 224 can be used to characterize the user state of the computing device 1101, the sensors 1121 of the computing device 1101, and the in-vehicle computing device 110 NIt should be understood that data can be obtained based on any other data such as the sensor(s) of the vehicle 100A, the sensor(s) of the vehicle 100A, and / or first-party data(s) from first-party applications(s) 182, and / or third-party data(s) from third-party applications(s) 183.

[0029] Furthermore, for example, a vehicle context signal(s) 225 obtained via the user context engine 125 can characterize the state of vehicle 100A. The state of vehicle 100A may include, for example, whether vehicle 100A is powered and in the "on" state, whether vehicle 100A is moving, whether vehicle 100A is stopped, whether vehicle 100A is parked, whether vehicle 100A is following a specific route, or whether vehicle 100A includes additional occupants other than the user of vehicle 100A. The state of vehicle 100A may additionally or alternatively include information associated with the operation of vehicle 100A, such as the amount of energy sources available to vehicle 100A (e.g., gas, batteries, etc.), whether any dashboard indicators are lit, and / or any other signals that characterize the state of the vehicle. Thus, the vehicle context signal(s) 225 can characterize the state of vehicle 100A. N This can be obtained based on the sensor(s)(or more), the sensor(s)(or more) of the vehicle 100A, and / or any other data such as OEM data 281 from one or more of the OEM applications 181.

[0030] Furthermore, the OEM data 281 is accessible by the OEM application associated with the vehicle 100A and may include any information available to the search engine 127 and / or response engine 128 when retrieving a given candidate response 228 for queries contained in the oral utterances captured in the audio data 201. Similarly, the first-party data 282 and third-party data 283 are accessible by the first-party application(s) 182 and third-party application(s) 183, respectively, and may include any information available to the search engine 127 and / or response engine 128 when retrieving a given candidate response 228 for queries contained in the oral utterances captured in the audio data 201, such as user account data, application usage information, and / or other data. As used herein, the term “first-party application” may refer to a software application developed and / or maintained by the same entity that develops and / or maintains the automated assistant and / or bias system 120 described herein. Furthermore, as used herein, the term “third-party application” may refer to a software application or system developed and / or maintained by an entity different from the one that develops and / or maintains the automated assistant and / or bias system 120 described herein.

[0031] The bias engine 126 can process these various signals and data, and / or other signals and data, to determine whether to restrict the query's search space to one or more specific corpus data. In other words, the bias engine 126 can process these various signals and data to determine whether one or more bias criteria are met. In response to determining that one or more of the bias criteria are met, the bias engine 126 can generate bias data 226 to provide to the search engine 127 and modify how the response(s) are obtained. For example, as will be described in more detail below with respect to Figure 3, the search engine 127 can restrict a search performed based on a query embodied in at least the request data 222 to a given corpus data from among several corpora 127A. In this example, the search engine 127 can prevent any search from being performed against other corpus data. For example, content items 227 obtained in response to a search being performed against a given corpus data may contain information identified from the given corpus data. Furthermore, as will be explained in more detail below with respect to Figure 4, for example, the search engine 127 can perform multiple searches on multiple corpora 127A. Therefore, the content items 227 obtained in response to performing multiple searches on multiple corpora 127A may contain information identified from each of the multiple corpora 127A. However, in this example, the response engine 128 may then apply a bias to the queries provided for presentation to the user for the given corpus data, selecting a given response.

[0032] As described herein, multiple corpora 127A may include a first corpus data and, in addition to the first corpus data, at least a second corpus data. The first corpus data may correspond, for example, to user manual corpus data of multiple user manuals provided by the OEM of vehicle 100A and / or other vehicles. The at least second corpus data may correspond to any other corpus data. In particular, the first corpus data may be defined or indexed at various levels of granularity. For example, user manual corpus data may be defined or indexed by the OEM, further defined or indexed by the year the vehicle was manufactured by the OEM, further defined or indexed by the manufacturer of the vehicle manufactured by the OEM, further defined or indexed by the model of the vehicle manufactured by the OEM, and further defined or indexed by multiple feature tags of the features of the vehicle manufactured by the OEM. Multiple feature tags associated with any document in the first corpus data may define various features of the vehicle described in that document (e.g., a document describing power seats, a document describing heated seats, a document describing sport mode, etc.). Therefore, when an oral utterance is received from a user inside the vehicle 100A, the year of manufacture, manufacturer, model, and / or feature tags of the vehicle 100A can be transmitted to the bias system 120 used when searching at least the first corpus data (for example, as part of request data 222 determined on the basis of processing the oral utterance, and / or as part of bias data 226 determined on the basis of OEM data 281 received from the OEM application 181).For example, when a user is inside vehicle 100A and provides an oral utterance, the vehicle 100A's year of manufacture, manufacturer, model, and a set of feature tags can be provided to the bias system 120, which can then be used to restrict searches in the user manual corpus data to relevant documents (e.g., documents associated with the vehicle 100A's year of manufacture, manufacturer, and model, and documents whose feature tags are a subset of the feature tags associated with vehicle 100A).

[0033] In various embodiments, the first corpus data may additionally or alternatively be associated with a number of notable terms that exist throughout the entire user manual corpus data and can be identified using various techniques (e.g., TF-IDF and / or other techniques across the entire user manuals contained in the user manual corpus data). These notable terms can be stored in association with various user manuals in the user manual corpus data stored in multiple corpora 127A. For example, a first user manual in the user manual corpus data may include a first set of notable terms, and a second user manual in the user manual corpus data may include a second set of notable terms. In particular, some of these notable terms may overlap between different user manuals contained in the user manual corpus data. In these embodiments, the request data 222 may include an indication that one or more terms of an oral utterance (e.g., determined based on the NLU output generated when processing the oral utterance) match one or more of the notable terms in the user manual corpus data using various techniques (e.g., soft matching). In these embodiments, the presence of one or more terms in an oral utterance that match one or more prominent terms in the user manual corpus data can be used as a bias criterion when biasing searches performed on the first and / or second corpus data. In some of these embodiments, when an oral utterance is received from a user inside the vehicle 100A, one or more terms in the oral utterance that match one or more prominent terms in the owner manual corpus of the vehicle 100A can be sent to a bias system 120 used when searching at least the first corpus data (e.g., as part of request data 222 determined on the basis of processing the oral utterance, and / or as part of bias data 226 determined on the basis of OEM data 281 received from OEM application 181).

[0034] The response engine 128 can analyze content item(s) 227 to identify a given candidate response(s) to be presented to the user in response to a query contained in the spoken utterance captured in the audio data 201. In some embodiments, the response engine 128 can extend content item(s) 227 with other content stored in the response(s) database 128A. For example, suppose a given content item corresponds to "tire pressure = 32 psi". In this example, the auto assistant can estimate that the current tire pressure of the passenger-side tire is 28 psi based on vehicle context signals(s) 225. Assuming the query contained in the oral utterance relates to tire pressure and asks why a light associated with tire pressure is illuminated (for example, biasing the user manual corpus data based on at least one or more vehicle context signals 225), the response engine 128 can expand a given content item to generate a given candidate response 228, such as "The tire pressure is 28 psi, but it should be 32 psi." Thus, the given candidate response 228 can, in response to the query, use context-relevant signals to provide the user with auditory and / or visual presentations and bias the interpretation of the oral utterance.

[0035] The example in Figure 2 illustrates queries contained in oral utterances captured in audio data 201A, but please understand that this is for illustrative purposes only and not intended to be limiting. For example, input processing engine 113 1-N / 122 can additionally or alternatively process non-audio data 201B that captures the query. For example, the user can additionally or alternatively process non-audio 201B, including text data, touch data, and / or any other non-audio data, in order to identify the query contained in non-audio data 201B. Thus, even when the user provides the query via non-audio data, the bias system 120 can use the same or similar techniques to identify and provide a given candidate response for presentation to the vehicle user.

[0036] Referring here to Figure 3, a flowchart is shown illustrating an exemplary method 300 for biasing the interpretation of oral utterances received in a vehicle environment. For convenience, the operation of method 300 is described with reference to the system that performs the operation. This system of method 300 comprises at least one processor, at least one memory, and / or other components of a computing device (for example, computing device 110 in Figure 1). 1-N This includes the bias system 120 in Figure 1, the computing device 710 in Figure 7, the remote server(s), and / or other computing devices. The operations of Method 300 are shown in a specific order, but this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0037] In block 352, the system receives an oral utterance from the user via a computing device, the oral utterance being provided while the user is inside the user's vehicle. The computing device may be, for example, the vehicle's onboard computing device, or the user's mobile computing device in the vehicle, which is communicably connected to the vehicle and / or the vehicle's onboard computing device. In some embodiments, the oral utterance may be received in response to the automated assistant being explicitly invoked on the computing device via a specific word or phrase (e.g., “Assistant,” “Hey Assistant,” etc.), operation of a hardware or software button on the computing device, a specific gesture or combination of gestures (e.g., hand movements, eye movements, gaze, and / or any one or a combination thereof), and / or other techniques for explicitly invoking the automated assistant. In additional or alternative embodiments, the oral utterance may be received in response to the automated assistant being implicitly invoked on the computing device, for example, based on one or more contextual signals associated with the user and / or the user's vehicle.

[0038] For example, oral utterances containing queries can be captured as audio data generated by one or more microphones on a computing device, which can optionally respond to an automated assistant call as described above. The system can use an automatic speech recognition (ASR) model to process the audio data capturing the oral utterances containing queries and generate ASR data for the queries, including one or more recognized terms corresponding to the queries. Furthermore, the system can use a natural language understanding (NLU) model to process the ASR data for the queries and generate NLU data for the queries, including intents, slot values ​​for parameters associated with those intents, and / or other NLU data. Based on the NLU data for the queries, the system can determine that the oral utterances contain queries. Furthermore, the system can utilize the NLU data when performing searches against one or more corpus datasets, as described below.

[0039] In block 354, the system obtains a corresponding vehicle sensor data instance of vehicle sensor data, which is generated by one or more vehicle sensors of the user's vehicle. One or more vehicle sensors of the user's vehicle may include, for example, a vehicle tire pressure sensor(s) that generates tire pressure data for the vehicle's tires, a vehicle airflow sensor(s) that generates airflow data for the vehicle's air conditioning system, a vehicle speed sensor(s) that generates vehicle speed data for the vehicle, an energy sensor(s) that generates energy source data for the vehicle's energy source, a transmission sensor(s) that generates transmission data for the vehicle's transmission, and / or any other sensors integrated into the vehicle and / or in-vehicle computing device. In particular, each of these sensors may generate a corresponding vehicle sensor data instance at a corresponding frequency, which may be the same for one or more vehicle sensors and different for one or more other sensors of the vehicle. In some implementations, a corresponding vehicle sensor data instance can be associated with a corresponding timestamp, which corresponds to the time the corresponding vehicle sensor data instance was generated by one or more of the vehicle sensors and / or the time the corresponding vehicle sensor data instance was acquired by the system. This time can be global time relative to a global clock (e.g., generated and / or acquired at 9:49 a.m.) or relative time relative to a relative clock (e.g., generated and / or acquired 2 minutes after the user enters the vehicle).

[0040] In block 356, the system determines whether to perform a first search in a first corpus data and / or a second search in a second corpus data to identify one or more candidate responses to the query contained in the oral utterance, based on (i) a query and (ii) a corresponding vehicle sensor data instance. In other words, the system may determine whether to bias the search space used to identify one or more candidate responses by potentially restricting the search space based on the query and the corresponding vehicle sensor data instance. The first corpus data may, for example, correspond to user manual corpus data provided by the vehicle's original equipment manufacturer (OEM), which is vehicle-specific. Furthermore, the second corpus data may, for example, correspond to corpus data available to the system that is not vehicle-specific (e.g., not provided by the vehicle's OEM), which could be, for example, web-based corpus data, a knowledge graph, and / or any other non-vehicle-specific corpus data available to the system. Furthermore, the system may make this bias determination based on one or more bias criteria.

[0041] In some embodiments, the system may make this bias determination based on determining whether a verbal utterance containing a query was received within a threshold duration for which a corresponding vehicle sensor data instance is generated, and / or as indicated by a corresponding timestamp associated with the corresponding vehicle sensor data instance. The threshold duration may include, for example, a static duration (e.g., 30 seconds, 2 minutes, etc.) or a dynamic duration based on other factors. For example, suppose the corresponding vehicle sensor data instance corresponds to a tire pressure data instance indicating that the passenger-side tire is low in air. As a result, an indication of low tire pressure may be provided to the user, such as the vehicle's dashboard light illuminating and / or an alert being displayed on the vehicle's onboard computing system display. Furthermore, suppose the system receives the verbal utterance, "Which one has low pressure?", within 15 seconds after the dashboard light illuminates and / or the alert is displayed. In this example, the system may determine that the verbal utterance contains a query associated with the dashboard light illuminating and / or the alert being displayed. Therefore, the system can restrict the search space for this query to a first corpus data corresponding to vehicle-specific user manual corpus data when retrieving response content in response to the query. If there is no such temporal relationship between the spoken utterance and the corresponding vehicle sensor data instance, the system can search not only the vehicle-specific user manual corpus data but also one or more other non-vehicle-specific corpus data when retrieving response content in response to the query. In other words, the system can use this temporal relationship between the spoken utterance and the corresponding vehicle sensor data instance to infer that the user provided the spoken utterance to further inquire about the illuminated dashboard lights and / or displayed alerts.

[0042] In additional or alternative embodiments, the system may make this bias determination based on determining whether a spoken utterance containing a query is related to a corresponding vehicle sensor data instance. For example, again, suppose the corresponding vehicle sensor data instance corresponds to a tire pressure data instance indicating that the passenger-side tire is low in air, the vehicle's dashboard light is on and / or an alert is displayed, and the system receives the spoken utterance, "Which one has low pressure?". As described above, the system can use an ASR model to process the audio data capturing the spoken utterance containing the query to generate one or more recognized terms corresponding to the query. This allows the system to determine, based on the corresponding vehicle sensors, that the vehicle's dashboard light is on and / or an alert is displayed, that the recognized term "low pressure" is related to the passenger-side tire being low in air. In this case, the system can utilize various word matching techniques (e.g., soft word matching, semantic word matching, and / or other techniques) to determine that the alert and the query are related. Therefore, the system can restrict the search space for this query to a first corpus data corresponding to vehicle-specific user manual corpus data when retrieving response content in response to the query. If there is no linguistic relationship between the query terms and the corresponding vehicle sensor data instances, the system can search not only vehicle-specific user manual corpus data but also one or more other non-vehicle-specific corpus data when retrieving response content in response to the query. In other words, the system can additionally or alternatively infer that the user provided the utterance to further inquire about the illuminated dashboard lights and / or displayed alerts, utilizing this linguistic relationship between the utterance and the corresponding vehicle sensor data instances.

[0043] In additional or alternative embodiments, the system may make this bias determination based on determining whether the user who provided the oral utterance was associated with the vehicle for a threshold duration. The threshold duration may represent, for example, the amount of time the user spent in the vehicle, the distance the user drove the vehicle, the number of times the user started the vehicle, the number of times a given corresponding vehicle sensor data instance generated by one or more sensors of the vehicle was captured, and / or other factors. For example, when processing audio data that captures oral utterances containing a query, the system may perform speaker identification on the audio data to determine whether the user who provided the oral utterance is a known user. The system may perform speaker identification using any known technique (e.g., text-dependent speaker identification, text-independent speaker identification, etc.) and any known speaker identification model. The system may additionally or alternatively use other technologies, such as facial recognition based on processing of visual data generated by a computing device or additional computing device's visual sensor(s), fingerprint recognition based on processing of fingerprint data generated by a computing device or additional computing device's fingerprint sensor(s), and / or any other technology, to determine whether the user providing the spoken utterance is a known user.

[0044] Furthermore, the system can compare the identification information of a known user of the vehicle (e.g., an account associated with the vehicle or auto assistant) with the identification information to determine whether the user is a known user. In these embodiments, if the user is a known user but is not associated with the vehicle for a threshold duration, the system may restrict the search space for this query to a first corpus data corresponding to vehicle-specific user manual corpus data when retrieving response content in response to the query. Additionally or alternatively, if the user is not a known user, the system may also restrict the search space for this query to a first corpus data corresponding to vehicle-specific user manual corpus data when retrieving response content in response to the query. In other words, the system may additionally or alternatively restrict the search space to user manual corpus data based on the fact that the user who provided the oral utterance is a new owner (or renter) of the vehicle who may not be familiar with the cause of the vehicle's dashboard lighting up and / or the cause of the alert being displayed. However, in these embodiments, if the user is a known user and associated with the vehicle for a threshold duration, the user is likely already well aware of what causes the vehicle's dashboard to light up and / or what causes the alert to appear, so the system does not need to restrict the search space for this query to a first corpus data corresponding to vehicle-specific user manual corpus data when retrieving response content in response to the query.

[0045] In additional or alternative embodiments, the system may make this bias determination based on determining whether a spoken utterance contains an explicit instruction to search only the first corpus data. For example, again, suppose a corresponding vehicle sensor data instance corresponds to a tire pressure data instance indicating that the passenger-side tire is low in air, and the vehicle's dashboard light illuminates and / or displays an alert. However, suppose the system receives the spoken utterance, “Check the user manual to see which one is low in air.” In this example, the query contained in the spoken utterance also includes an explicit instruction to search only the first corpus data, as indicated in “user manual.”

[0046] In additional or alternative embodiments, the system may make this bias decision based on whether the oral utterance was received while the user was inside the vehicle. For example, suppose the oral utterance "Which one has low air pressure?" was received by the system while the user was inside the vehicle. In this example, even if the system has decided to perform both a first search on the first corpus data and a second search on the second corpus data (as described, for example, in the embodiment of Figure 4), it may bias the selection of a given candidate response towards one or more of the candidate responses retrieved from the first corpus data based on the oral utterance received while the user was inside the vehicle.

[0047] In additional or alternative embodiments, the system may make this bias determination based on whether the oral utterance contains one or more terms that have been determined to correspond to one or more prominent terms. For example, suppose the oral utterance "What is preconditioning?" is received by the system while the user is inside a vehicle. In this example, the NLU output generated based on the processing of the oral utterance may contain the term "preconditioning". Furthermore, the system cross-references the term "preconditioning" with one or more prominent terms in the user manual corpus data and determines (e.g., using various word matching techniques) that "preconditioning" is a prominent term associated with one or more documents in the first corpus data. As a result, the system performs a first lookup on the first corpus data, but does not have to perform one on the second corpus data (e.g., as described with respect to Figure 3). Additionally or alternatively, the system may perform a second search on a second corpus data (as described, for example, in the embodiment of Figure 4), and may bias the selection of a given candidate response towards one or more candidate responses obtained from the first corpus data, based on oral utterances containing one or more prominent terms from the user manual corpus data.

[0048] If, in a repetition of block 356, the system decides to perform a first search on the first corpus data but not a second search on the second corpus data, the system proceeds to block 358. In block 358, the system allows the first search to be performed on the first corpus data to identify one or more candidate responses to the query. In block 360, the system prevents the system from performing a second search on the second corpus data to identify one or more of the candidate responses to the query. In other words, the system can generate a first search to submit to the first corpus data based on the query's NLU data and corresponding vehicle sensor data instances, but can refrain from generating any second search to submit to the second corpus data. For example, suppose again the system receives the oral utterance, "What's under pressure?", and the system decides to restrict the search space to only the user manual corpus data. In this example, the system can generate a structured request to search the user manual corpus data based on the term “under pressure” and an indication that a tire pressure data instance indicates low air pressure on the passenger side. In response, based on one or more content items contained in the user manual corpus data, one or more candidate responses can be identified, such as what the tire pressure should be, what tire pressure will cause the vehicle's dashboard lights to turn on and / or display an alert, and / or any other content items contained in the user manual corpus data retrieved in response to the first search. In this example, the search could be further limited to the “Tires” section of the user manual corpus data.

[0049] If, in an iteration of block 356, the system decides to perform a first search on the first corpus data and a second search on the second corpus data, the system proceeds to block 362. In block 362, the system performs a first search on the first corpus data to identify one or more candidate responses to the query. In block 364, the system performs a second search on the second corpus data to identify one or more of the candidate responses to the query. In other words, the system can generate a first search to submit to the first corpus data based on the query's NLU data and corresponding vehicle sensor data instances, and can also generate at least a second search to submit to the second corpus data based on the query. For example, the system may perform a first search on the first corpus data to identify one or more candidate responses to the query in the same or similar manner as described above with respect to the operation of block 358 and the user manual corpus data. Furthermore, the system also performs a second search on a second corpus of data to identify one or more candidate responses to a query, such as queries submitted to a search engine and / or one or more applications. For example, a query could be submitted to a media application to retrieve response content corresponding to the song "Under Pressure" by the band Queen.

[0050] In block 366, the system provides a given candidate response to the user from among one or more candidate responses via a computing device or an additional computing device. In some embodiments, the one or more candidate responses may include only candidate responses based on a first corpus data (e.g., in embodiments where the system reaches block 360 to block 366). In other embodiments, the one or more candidate responses may include candidate responses based on both the first and second corpus data (e.g., in embodiments where the system reaches block 364 to block 366). In some embodiments, the system can select a given candidate response from among one or more candidate responses based on a ranking of the one or more candidate responses. The system can rank one or more candidate responses based on one or more ranking criteria. One or more ranking criteria may include, for example, one or more terms in a query, corresponding vehicle sensor data instances, one or more user context signals characterizing the state of a vehicle user, one or more vehicle context signals characterizing the state of a vehicle, application data associated with one or more applications accessible by a computing device or an additional computing device, and / or other criteria. Furthermore, the system can select a given candidate response based on ranking and provide the given candidate response to the user audibly and / or visually via a computing device or additional computing devices.

[0051] Referring here to Figure 4, a flowchart is shown illustrating another exemplary method 400 for biasing the interpretation of oral utterances received in a vehicle environment. For convenience, the operation of method 400 is described with reference to the system that performs the operation. This system of method 400 comprises at least one processor, at least one memory, and / or other components of a computing device (for example, computing device 110 in Figure 1). 1-N This includes the bias system 120 in Figure 1, the computing device 710 in Figure 7, the remote server(s), and / or other computing devices. The operations of Method 400 are shown in a specific order, but this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0052] In block 452, the system receives a verbal utterance containing a query from the user via a computing device, which is provided while the user is inside the user's vehicle. In block 454, the system obtains a corresponding vehicle sensor data instance of vehicle sensor data, which is generated by one or more vehicle sensors in the user's vehicle. The operations of blocks 452 and 454 of method 400 in Figure 4 can be performed in the same or similar manner as described above with respect to the operations of blocks 352 and 354 of method 300 in Figure 3.

[0053] In block 456, the system processes the oral utterance to identify one or more candidate responses to a query contained in the oral utterance. For example, as shown in block 456A, the system performs a first lookup on a first corpus data to identify one or more first candidate responses to the query. The system can perform a first lookup on a first corpus data to identify one or more first candidate responses to the query in the same or similar manner as described above with respect to the operation of blocks 358 and 360 of method 300 in Figure 3. Furthermore, as shown in block 456B, the system can perform a second lookup on a second corpus data to identify one or more second candidate responses to the query. The system can perform a second lookup on a second corpus data to identify one or more second candidate responses to the query in the same or similar manner as described above with respect to the operation of block 364 of method 300 in Figure 3. In particular, in contrast to method 300 in Figure 3, method 400 in Figure 4 does not require a decision to restrict the search space before performing any of the searches.

[0054] Rather, in block 458, the system determines whether to bias one or more of the first candidate responses identified based on a first search of the first corpus data into one or more of the first candidate responses. In other words, the system obtains one or more first candidate responses to the query based on a first search of the first corpus data and one or more second candidate responses to the query based on a second search of the second corpus data, and then determines whether to bias one or more of the first candidate responses based on one or more bias criteria. The one or more bias criteria used in making this decision are described in more detail herein (for example, relating to the operation of block 356 of method 300 in Figure 3).

[0055] In an iteration of block 458, if the system decides to bias one or more of the candidate responses to one or more of the first candidate responses identified based on a first search of the first corpus data, the system proceeds to block 460. In block 460, the system selects one of the one or more first candidate responses from among the one or more candidate responses as a given candidate response to be presented to the user. In embodiments where one or more first candidate responses include multiple candidate responses, the system may rank the one or more first candidate responses and select, for example, the highest-ranked first candidate response as a given candidate response. The system may rank one or more first candidate responses using various ranking criteria described herein.

[0056] In a repeat of block 458, if the system decides not to bias the selection of one or more candidate responses from the first candidate responses identified based on a first search of the first corpus data, the system proceeds to block 462. In block 462, the system selects from the one or more candidate responses one or more first candidate responses or one or more second candidate responses as a given candidate response to be presented to the user. Similarly, the system may rank the one or more first candidate responses and one or more second responses and select, for example, the highest-ranked one or more first candidate responses and one or more second candidate responses as a given candidate response. The system may rank the one or more first candidate responses using various ranking criteria described herein.

[0057] In block 464, the system is made to provide a given candidate response to the user via a computing device or an additional computing device. The system is made to provide a given candidate response to the user in the same or similar manner as described above with respect to the operation of block 366 of method 300 in Figure 3.

[0058] Referring here to Figure 5, a flowchart is shown illustrating another exemplary method 500 for biasing the interpretation of oral utterances received in a vehicle environment. For convenience, the operation of method 500 is described with reference to the system that performs the operation. This system of method 500 comprises at least one processor, at least one memory, and / or other components of a computing device (for example, computing device 110 in Figure 1). 1-N This includes the bias system 120 in Figure 1, the computing device 710 in Figure 7, the remote server(s), and / or other computing devices. The operations of Method 500 are shown in a specific order, but this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0059] In block 552, the system receives a request from the user to access user manual corpus data specific to the user's vehicle via a computing device. In block 554, the system accesses user manual corpus data provided by the vehicle's original equipment manufacturer (OEM). The computing device can be, for example, the user's mobile computing device, and the request can target an automated assistant running at least partially on the user's mobile computing device. In some implementations, the request can be contained in text or touch data provided by the user on the computing device's display. For example, the system may receive a request based on the user providing text or touch input detected through the interface of an automated assistant application. In other embodiments, the request can be contained in an oral utterance received in response to an explicit or implicit invocation of the automated assistant, as described in relation to the operation of block 352 of Method 300 in Figure 3.

[0060] For example, a spoken utterance containing a request can be captured as audio data generated by one or more microphones on a computing device, which can optionally respond to an automated assistant call as described above. The system can use an automatic speech recognition (ASR) model to process the audio data capturing the spoken utterance containing the request and generate ASR data for the request, including one or more recognized terms corresponding to the request. Furthermore, the system can use a natural language understanding (NLU) model to process the ASR data for the request (or text or touch data provided by the user) and generate NLU data for the request, including intent(s), slot values(s) for parameters(s) associated with the intent(s), and / or other NLU data. Based on the NLU data for the request, the system can determine that the spoken utterance contains a request to access user manual corpus data specific to the user's vehicle. Furthermore, the system can generate the request and send it to a third-party application or agent associated with the vehicle's OEM. The request may include, for example, the manufacturer and model of the user's vehicle, the year of manufacture associated with the manufacturer and model of the user's vehicle, and / or other information that can be used to identify the correct user manual corpus data for the user's vehicle. In this example, the system may access this information based on user profile data from the user's user profile and / or prompt the user to provide this information. Furthermore, the system may access the user manual corpus data from a third-party application or agent associated with the vehicle's OEM and responding to the submission of the request.

[0061] In block 556, the system receives an oral utterance from the user via a computing device, which includes a query targeting user manual corpus data. In some embodiments, the oral utterance may be received in response to the automatic assistant being explicitly or implicitly invoked on the computing device, as described above with respect to the operation of block 552. Furthermore, the system can process the audio data capturing the oral utterance using at least the ASR model and / or NLU model to determine that the oral utterance includes a query targeting user manual corpus data. In particular, in various embodiments, the request received in block 552 and the oral utterance received in block 556 may be received from the user as a single request.

[0062] In block 558, the system performs a search against at least the user manual corpus data to identify one or more candidate responses to a query contained in a spoken utterance, and this search against at least the user manual corpus data is query-based and does not utilize any corresponding vehicle sensor data instance of vehicle sensor data generated by one or more sensors of the vehicle. For example, suppose the user's vehicle is from 1994 or later. In this example, the vehicle is unlikely to have an on-board computing device that can communicate with the system. As a result, the system may not have access to a corresponding sensor data instance when performing a search against the user manual corpus data to identify one or more candidate responses to a query. Nevertheless, an automated assistant running on the user's mobile computing device may still be able to access the user manual corpus data for this vehicle, allowing the user to query the user manual corpus data via the automated assistant.

[0063] In block 560, the system is made to provide a given candidate response to the user from among one or more candidate responses via a computing device. In some embodiments, the system can select a given candidate response from among one or more candidate responses based on a ranking of the candidate responses. The system can rank one or more candidate responses based on one or more ranking criteria. One or more ranking criteria may include, for example, one or more terms in a query, corresponding vehicle sensor data instances, one or more user context signals characterizing the state of a vehicle user, one or more vehicle context signals characterizing the state of a vehicle, application data associated with one or more applications accessible by the computing device or an additional computing device, and / or other criteria. Furthermore, the system can select a given candidate response based on the ranking and provide the given candidate response to the user audibly and / or visually via the computing device or an additional computing device.

[0064] While Method 500 in Figure 5 describes performing a search on user manual corpus data only, it should be understood that this is for illustrative purposes only and not intended as an limitation. For example, the system can perform additional searches on additional corpus data that is not specific to vehicles. In these embodiments, a given candidate response can be further selected from any other candidate responses determined based on the search on the additional corpus data.

[0065] Referring here to Figures 6A and 6B, various non-exclusive examples of computing devices that demonstrate various user interactions that bias the speech processing of oral utterances in a vehicle environment are shown. Specifically referring to Figure 6A, the in-vehicle computing device 110 from Figure 1 N This is shown. In-vehicle computing device 110N The display 620 has multiple different parts targeting different applications. N This includes, for example, display 620 in Figure 6A. N This is at least partially the in-vehicle computing device 110 N Display 620 for automated assistant applications running on N Part 1, 622 N And, at least partially, the in-vehicle computing device 110 N Display 620 for OEM applications of vehicle 100A running on OEM N Part 2, 624 N And, display 620 targeting third-party media applications associated with exemplary music streaming services N Part 3, 626 N Includes. In-vehicle computing device 110 N 620 display N However, although it is shown in Figure 6A as having a specific configuration (for example, multiple different parts targeting various applications), please understand that this is for illustrative purposes only and is not intended to be limiting. For example, display 620 N It can be configured in any preferred manner and may be specific to a particular OEM.

[0066] For example, suppose that a corresponding vehicle sensor data instance for a tire pressure data instance is generated by one or more tire sensors on vehicle 100A, indicating that the tire pressure of one or more tires on vehicle 100A is low. Furthermore, display 620 targeting OEM applications N Part 2, 624 N However, based on the tire pressure data instance, alert 624A appears saying "Tire pressure is low". NLet's assume that this was displayed to present to the user. Furthermore, the user of vehicle 100A called the automatic assistant and asked query 622A, "What's under pressure?" N Assume that you have provided an oral utterance that includes . In this example, we process the audio data that captured the oral utterance so that the oral utterance matches query 622A N It can be decided to include it.

[0067] Furthermore, using the various biasing techniques described herein (for example, with respect to Figures 3 and 4), the automated assistant can perform query 622A. N When providing alert 624A, the user N It can be determined that the user is requesting an explanation regarding alert 624A. In particular, the automated assistant can determine that the user is not submitting a general query, but rather alert 624A. N For example, query 622A is requesting an explanation regarding this matter. N and alert 624A N The temporal relationship between (for example, alert 624A N Query 622A is provided to the user for presentation within a threshold duration. N (Based on the receipt of) Query 622A N and alert 624A N The linguistic relationship between (for example, query 622A) N and alert 624A N The determination can be based on the fact that both include the term "under pressure", the temporal relationship between the user and vehicle 100A (e.g., based on the length of time the user has owned vehicle 100A, based on the distance the user has driven vehicle 100A, based on the number of times the user has started the vehicle, etc.), and / or other bias criteria.

[0068] Therefore, in the example in Figure 6A, the automated assistant queries query 622A. NThe search space for identifying one or more candidate responses to can be limited to the user manual corpus data, as described in Figure 3. Furthermore, given candidate response 622B, "The tire pressure is 28 psi, but according to the user manual it should be 32 psi." N This may be provided to the user visually (as shown in Figure 6A) and / or audibly. Additionally or alternatively, in the example in Figure 6A, the automated assistant provides query 622A N To identify one or more candidate responses to 622B, both the user manual corpus data and additional corpus data (e.g., the internet, other applications, other databases, etc.) can be searched, but as described in Figure 4, one or more given candidate responses 622B N The selection can be biased towards one or more candidate responses obtained using user manual corpus data. In these additional or alternative examples, the automated assistant can also handle query 622A. N This may be related to 622A N It is also possible to identify other candidate responses that are not considered to be a response to the above. For example, in various embodiments, the automated assistant may receive a notification 626A that says, "Click here to play Under Pressure by Queen." N Display 620 for third-party media applications N Part 3, 626 N Further information can be presented to the user through this method.

[0069] Referring specifically to Figure 6B, the in-vehicle computing device 1101 from Figure 1 is shown. While the computing device 1101 is shown as a mobile computing device for the user of vehicle 100A, it should be understood that this is for illustrative purposes only and not intended as an limitation. The computing device 1101 includes a display 6201 having various system interface elements 6811, 6821, and 6831 (e.g., hardware and / or software interface elements) that can be interacted with by the user and cause the computing device 1101 to perform one or more actions. The display 6201 of the computing device 1101 allows the user to interact with the content displayed on the display 6201 by typing or touch input (for example, by directing user input to the display 6201 or the text interface element 6841) and / or by oral input (for example, by selecting the microphone interface element 6851 on the computing device 1101, or simply by speaking without necessarily selecting the microphone interface element 6851 (i.e., the automated assistant may monitor one or more terms or phrases, gestures, gaze, mouth movements, lip movements, and / or other conditions that enable oral input)). Furthermore, an automated assistant application can be implemented on the computing device 1101, at least in part, as indicated by 6221.

[0070] For example, suppose the user of vehicle 100A has already provided the automated assistant with a request to access user manual corpus data associated with vehicle 100A. Furthermore, suppose the user has provided input 622A1, “What is preconditioning?” (e.g., via verbal utterance or typed input). In this example, the automated assistant can perform a lookup against the user manual corpus data provided by the vehicle's OEM to identify a given candidate response 622B1, “In the context of the vehicle, preconditioning allows the vehicle to be preheated or precooled before entering the vehicle’s cabin,” without using any vehicle sensor data instance of the vehicle sensor data. Without the techniques described herein, the automated assistant may not return any definition in the context of vehicle 100A. Nevertheless, other definitions of “preconditioning” may be provided to the user, as indicated by a notice 622C1, “Click here for other definitions of preconditioning.” In particular, the embodiment shown in Figure 6B is especially advantageous when there is no in-vehicle computing device and / or when the in-vehicle computing device 110N is unable to perform the Auto Assistant, as described with respect to Figure 6A. Please understand that while Figures 6A and 6B illustrate specific embodiments, they are not intended to be limiting.

[0071] Referring now to Figure 7, a block diagram of an exemplary computing device 710 that may be optionally used to perform one or more embodiments of the techniques described herein is shown. In some embodiments, one or more computing devices, one or more vehicles, and / or other components may constitute one or more components of the exemplary computing device 710.

[0072] The computing device 710 typically includes at least one processor 714 that communicates with numerous peripheral devices via a bus subsystem 712. These peripheral devices may include, for example, a storage subsystem 724 including a memory subsystem 725 and a file storage subsystem 726, a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices enable user interaction with the computing device 710. The network interface subsystem 716 provides an interface to an external network and connects to a corresponding interface device in another computing device.

[0073] The user interface input device 722 may include pointing devices such as keyboards, mice, trackballs, touchpads, and graphic tablets; audio input devices such as scanners, touchscreens integrated into displays, and speech recognition systems; microphones; and / or other types of input devices. In general, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computing device 710 or a communication network.

[0074] The user interface output device 720 may include non-visual displays such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include flat panel devices such as cathode ray tubes (CRTs) or liquid crystal displays (LCDs), projection devices, or other mechanisms for creating visible images. The display subsystem may also provide non-visual displays via audio output devices, etc. In general, the use of the term “output device” is intended to include all possible types of devices and methods for outputting information from the computing device 710 to a user or another machine or computing device.

[0075] The storage subsystem 724 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 724 may include logic that performs selected embodiments of the methods disclosed herein, and logic that implements the various components shown in Figures 1 and 2.

[0076] These software modules typically run on processor 714 alone or in combination with other processors. The memory 725 used by the storage subsystem 724 may include a number of memories, such as main random access memory (RAM) 730 for storing instructions and data during program execution, and read-only memory (ROM) 732 for storing fixed instructions. The file storage subsystem 726 can provide persistent storage of program files and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment may be stored by the file storage subsystem 726 within the storage subsystem 724, or on other machines accessible by the processor 714(or more).

[0077] The bus subsystem 712 provides a mechanism that enables various components and subsystems of the computing device 710 to communicate with each other as intended. Although the bus subsystem 712 is schematically shown as a single bus, alternative embodiments of the bus subsystem 712 may use multiple buses.

[0078] The computing device 710 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing systems or computing devices. Because computers and networks are constantly changing, the description of the computing device 710 shown in Figure 7 is intended only as a specific example to illustrate several embodiments. Many other configurations of the computing device 710 may have more or fewer components than the computing device shown in Figure 7.

[0079] Wherever the systems described herein may collect or monitor personal information relating to a user, or may use personal and / or monitoring information, the user may be provided with the opportunity to control whether the program or function collects user information (e.g., information relating to the user's social networks, social behavior or activities, occupation, user preferences, or current geographical location), or whether and / or how it receives content that may be more relevant to the user from a content server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identity may be processed so that personally identifiable information cannot be determined, or if geographical location information is obtained (to the level of city, zip code, or state, for example), the user's geographical location may be generalized so that the user's specific geographical location cannot be determined. Thus, the user may control how information relating to them is collected and / or used.

[0080] In some embodiments, a method is provided which is implemented by one or more processors, the method of receiving an oral utterance containing a query from a user via a computing device, the oral utterance being provided while the user is located inside the user's vehicle, and obtaining a corresponding vehicle sensor data instance of vehicle sensor data, the corresponding vehicle sensor data instance being generated by one or more vehicle sensors of the user's vehicle, and to identify one or more candidate responses to the query contained in the oral utterance, (i) the query and / or (ii) a first corpus data based on the corresponding vehicle sensor data instance. The process includes determining whether to perform a search and / or a second search on a second corpus of data, and, in response to a decision to perform a first search on the first corpus of data but not on the second corpus of data, causing the first search to be performed on the first corpus of data to identify one or more candidate responses to a query contained in an oral utterance, wherein the first search is based on (i) a query and (ii) a corresponding vehicle sensor data instance, and providing a given candidate response from among one or more candidate responses for presentation to the user via a computing device or additional computing device.

[0081] These and other embodiments disclosed herein may optionally include one or more of the following features:

[0082] In some embodiments, the method may further include using an automatic speech recognition (ASR) model to process audio data capturing oral utterances containing a query to generate ASR data for the query, and using a natural language understanding (NLU) model to process the ASR data to generate NLU data for the query. In some versions of these embodiments, performing a first search on a first corpus of data to identify one or more candidate responses for a query contained in an oral utterance may include performing a first search on the first corpus data to submit NLU data for the query and indications of corresponding vehicle sensor data instances to the first corpus data, and identifying one or more candidate responses based on the content responding to the NLU data for the query.

[0083] In some embodiments, the first corpus data may correspond to a vehicle-specific user manual corpus data provided by the vehicle's original equipment manufacturer (OEM), and the second corpus data may correspond to additional, non-vehicle-specific corpus data.

[0084] In some embodiments, the decision to perform a first search on first corpus data but not a second search on second corpus data may include identifying a corresponding timestamp associated with a corresponding vehicle sensor data instance, where the corresponding timestamp associated with the corresponding vehicle sensor data instance corresponds to the time the corresponding vehicle sensor data instance was generated; determining, based on the corresponding timestamp associated with the corresponding vehicle sensor data instance, whether an oral utterance containing a query was received within a threshold duration with respect to the time the corresponding vehicle sensor data instance was generated; and, in response to the determination that the oral utterance containing the query was received within a threshold duration, deciding to perform a first search on first corpus data but not a second search on second corpus data. In some versions of these embodiments, the method may further include deciding to perform a first search on first corpus data and a second search on second corpus data in response to the determination that the oral utterance containing the query was received within a threshold duration. In some further versions of these embodiments, the method may further include, in response to a decision to perform a first search on first corpus data and a second search on second corpus data, performing a first search on first corpus data to identify one or more candidate responses to a query contained in an oral utterance, wherein the first search is based on (i) a query and (ii) a corresponding vehicle sensor data instance; and performing a second search on second corpus data to identify one or more candidate responses to a query contained in an oral utterance, wherein the second corpus data is added to the first corpus data, and the second search is based on (i) a query but not on (ii) a corresponding vehicle sensor data instance.In further versions of these embodiments, the method may further include selecting a given candidate response to be provided for presentation to a user via a computing device or additional computing device, based on a ranking of one or more candidate responses. In further versions of these embodiments, the ranking of one or more candidate responses is based on one or more terms of a query, corresponding vehicle sensor data instances, one or more user context signals characterizing the user state of the vehicle, one or more vehicle context signals characterizing the state of the vehicle, or application data associated with one or more applications accessible on the computing device or additional computing device.

[0085] In some embodiments, deciding to perform a first search on a first corpus data but not a second search on a second corpus data may include determining that one or more terms in a query contained in an oral utterance are associated with a corresponding vehicle sensor data instance. In some versions of these embodiments, once the corresponding vehicle sensor data instance is generated, a computing device or additional computing device may provide an indication of the corresponding vehicle sensor data instance for presentation to the user. In several more versions of these embodiments, determining that one or more terms in a query contained in an oral utterance are associated with a corresponding vehicle sensor data instance may include determining that one or more terms in the query contained in the oral utterance are subject to an indication of the corresponding vehicle sensor data instance provided for presentation to the user via a computing device or additional computing device.

[0086] In some embodiments, the method may further include determining the durations a user is associated with a vehicle based on processing an oral utterance containing a query. Determining the durations a user is associated with a vehicle may include having speaker identification performed based on processing the oral utterance to determine whether the user is a known user, and, in response to determining that the user is a known user, determining the durations a user is associated with a vehicle based on a user profile associated with the user. In some versions of these embodiments, deciding to perform a first search on a first corpus data but not a second search on a second corpus data may include determining whether the durations a user is associated with a vehicle meet a threshold duration, and, in response to determining that the durations a user is associated with a vehicle do not meet a threshold duration, deciding to perform a first search on the first corpus data but not a second search on the second corpus. Further versions of these embodiments may further include the method performing a first look on a first corpus data to identify one or more candidate responses to a query contained in an oral utterance based on (i) a query and (ii) a corresponding vehicle sensor data instance, in response to the user determining that the duration associated with a vehicle satisfies a threshold duration; and performing a second look on a second corpus data to identify one or more candidate responses to a query contained in an oral utterance based on (i) a query, wherein the second corpus data is added to the first corpus data. Further versions of these embodiments may further include selecting a given candidate response to be provided for presentation to the user via a computing device or additional computing device, based on a ranking of one or more candidate responses.

[0087] In some embodiments, methods are provided that are implemented by one or more processors, the method comprising receiving an oral utterance from a user via a computing device, the oral utterance being provided by the user while the user is located inside the user's vehicle, and obtaining a corresponding vehicle sensor data instance of vehicle sensor data, the corresponding vehicle sensor data instance being generated by one or more vehicle sensors of the user's vehicle, and processing the oral utterance to identify a plurality of candidate responses to the query contained in the oral utterance. Processing an oral utterance to identify multiple candidate responses to a query contained in the oral utterance includes performing a first search on a first corpus data to identify one or more first candidate responses from among multiple candidate responses to a query contained in the oral utterance, wherein the first search is based on (i) a query and (ii) a corresponding vehicle sensor data instance; and performing a second search on a second corpus data to identify one or more second candidate responses from among multiple candidate responses to a query contained in the oral utterance, wherein the second corpus data is added to the first corpus data, and the second search is based on (i) a query but not on (ii) a corresponding vehicle sensor data instance. The method further includes selecting a given candidate response from among multiple candidate responses and providing the given candidate response for presentation to a user via a computing device or additional computing device.

[0088] These and other embodiments disclosed herein may optionally include one or more of the following features:

[0089] In some embodiments, the method may further include using an automatic speech recognition (ASR) model to process audio data capturing oral utterances containing a query to generate ASR data for the query, and using a natural language understanding (NLU) model to process the ASR data to generate NLU data for the query. In some versions of these embodiments, performing a first search on a first corpus data to identify one or more first candidate responses for a query contained in an oral utterance may include performing a first search on the first corpus data to bring the first corpus data to submit indications of NLU data for the query and corresponding vehicle sensor data instances, and identifying one or more first candidate responses based on the first content responding to the NLU data for the query.

[0090] In some further versions of these embodiments, performing a second search on a second corpus data to identify one or more second candidate responses to a query contained in an oral utterance may include having the second corpus data submit the NLU data of the query, performing a second search on the second corpus data, and identifying one or more second candidate responses based on the second content that responds to the NLU data of the query.

[0091] In further additional or alternative versions of these embodiments, the first corpus data may correspond to vehicle-specific user manual corpus data provided by the vehicle's original equipment manufacturer (OEM), and the second corpus data may correspond to non-vehicle-specific web-based corpus data.

[0092] In further additional or alternative versions of these embodiments, selecting a given candidate response from a plurality of candidate responses may include ranking the plurality of candidate responses, biasing the ranking of the plurality of candidate responses with respect to one or more first candidate responses, and in the biased ranking of the plurality of candidate responses, selecting a given candidate response from the plurality of candidate responses. In further versions of these embodiments, the ranking of the plurality of candidate responses is based on one or more terms of a query, corresponding vehicle sensor data instances, one or more user context signals characterizing the state of a vehicle user, one or more vehicle context signals characterizing the state of a vehicle, and application data associated with one or more applications accessible on a computing device or additional computing device.

[0093] In some embodiments, a method is provided which is implemented by one or more processors, the method of receiving an oral utterance from a user via a computing device, the oral utterance being received while the user is located inside the user's vehicle, and obtaining a corresponding vehicle sensor data instance of vehicle sensor data, the corresponding vehicle sensor data instance being generated by one or more vehicle sensors of the user's vehicle, and processing the oral utterance being received, the method of determining the duration to which the user is associated with the vehicle, and determining one or more candidate responses to the query contained in the oral utterance, by (i) the query, (ii) the corresponding vehicle sensor data instance, and / or (iii) the user Based on whether the duration associated with the vehicle does not meet a temporal threshold, the system includes determining whether to perform a first search on the first corpus data and / or a second search on the second corpus data, and, in response to a decision to perform a first search on the first corpus data but not a second search on the second corpus data, causing the system to perform a first search on the first corpus data to identify one or more candidate responses to the query contained in the oral utterance based on (i) the query and (ii) the corresponding vehicle sensor data instance, and providing a given candidate response from among the one or more candidate responses for presentation to the user via a computing device or additional computing device.

[0094] In some embodiments, a method is provided which is implemented by one or more processors, the method of receiving an oral utterance from a user via a computing device, the oral utterance being provided by the user while the user is located in the user's vehicle, and receiving a corresponding vehicle sensor data instance of vehicle sensor data, the corresponding vehicle sensor data instance being generated by one or more vehicle sensors of the user's vehicle, and identifying a corresponding timestamp associated with the corresponding vehicle sensor data instance, the corresponding timestamp associated with the corresponding vehicle sensor data instance being generated, and identifying one or more candidate responses to the query contained in the oral utterance, (i) the query, ( ii) determining whether to perform a first search on a first corpus data and / or a second search on a second corpus data based on a corresponding vehicle sensor data instance and / or a corresponding timestamp associated with the corresponding vehicle sensor data instance; and, in response to a decision to perform a first search on the first corpus data but not a second search on the second corpus data, (i) causing the first search to be performed on the first corpus data to identify one or more candidate responses to a query contained in an oral utterance based on a query and / or (ii) the corresponding vehicle sensor data instance, and providing a given candidate response from among the one or more candidate responses for presentation to the user via a computing device or additional computing device.

[0095] In some embodiments, a method is provided which is implemented by one or more processors, the method comprising: receiving a request from a user via a computing device to access user manual corpus data specific to the user's vehicle; accessing user manual corpus data from the vehicle's original equipment manufacturer (OEM) in response to receiving a request to access user manual corpus data specific to the vehicle; receiving an oral utterance from a user via a computing device containing a query relating to the user manual corpus data; and performing a search on the user manual corpus data to identify one or more candidate responses to the query contained in the oral utterance, the search on the user manual corpus data being (i) based on the query but (ii) not utilizing any corresponding vehicle sensor data instance of vehicle sensor data generated by one or more vehicle sensors of the vehicle; and providing a given candidate response from one or more candidate responses for presentation to the user via the computing device or additional computing devices.

[0096] In some embodiments, methods are provided that are implemented by one or more processors, the method comprising receiving an oral utterance from a user via a computing device, the oral utterance being provided by the user while the user is located inside the user's vehicle, and obtaining a corresponding vehicle sensor data instance of vehicle sensor data, the corresponding vehicle sensor data instance being generated by one or more vehicle sensors of the user's vehicle, and processing the oral utterance to identify a plurality of candidate responses to the query contained in the oral utterance. Processing an oral utterance to identify multiple candidate responses to a query contained in the oral utterance includes performing a first search on a first corpus data to identify one or more first candidate responses from among multiple candidate responses to a query contained in the oral utterance, wherein the first search is at least (i) based on a query; and performing a second search on a second corpus data to identify one or more second candidate responses from among multiple candidate responses to a query contained in the oral utterance, wherein the second corpus data is added to the first corpus data, and the second search is (i) based on a query. The method further includes (ii) selecting a given candidate response from among multiple candidate responses based on a corresponding vehicle sensor data instance, wherein the selection is (ii) biased towards one or more first candidate responses based on a corresponding vehicle sensor data instance; and providing the given candidate response for presentation to a user via a computing device or additional computing device.

[0097] Furthermore, some embodiments include one or more processors in one or more computing devices (e.g., a central processing unit (CPU(or more)), a graphics processing unit (GPU(or more)), and / or a tensor processing unit (or more) (TPU(or more)), where one or more processors are operable to execute instructions stored in associated memory, and the instructions are configured to perform any of the methods described above. Some embodiments also include one or more non-temporary computer-readable storage media that store computer instructions executable by one or more processors to perform any of the methods described above. Some embodiments also include a computer program product that includes instructions executable by one or more processors to perform any of the methods described above.

[0098] It will be understood that all combinations of the above concepts and additional concepts described in more detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter disclosed herein.

Claims

1. A method implemented by one or more processors, Receiving an oral utterance from a user via a computing device, the oral utterance being provided while the user is inside the user's vehicle, The acquisition of a corresponding vehicle sensor data instance, wherein the corresponding vehicle sensor data instance is generated by one or more vehicle sensors of the user's vehicle, (i) determining whether to perform at least the first search of a first search in a first corpus data and a second search in a second corpus data in order to identify one or more candidate responses to the query contained in the oral utterance based on the query and / or (ii) the corresponding vehicle sensor data instance, Identifying a corresponding timestamp associated with the corresponding vehicle sensor data instance, wherein the corresponding timestamp associated with the corresponding vehicle sensor data instance corresponds to the time the corresponding vehicle sensor data instance was generated. Based on the corresponding timestamp associated with the corresponding vehicle sensor data instance, it is determined whether the verbal utterance including the query was received within a threshold duration with respect to the time the corresponding vehicle sensor data instance was generated. In response to the determination that the oral utterance containing the query was received within the threshold duration, The decision is made to perform the first search on the first corpus data, but not to perform the second search on the second corpus data. In response to the decision to perform the first search on the first corpus data but not the second search on the second corpus data, To identify one or more candidate responses to the query contained in the oral utterance, a first search is performed on the first corpus data, wherein the first search is based on (i) the query and (ii) the corresponding vehicle sensor data instance. This includes providing a given candidate response from among the one or more candidate responses to the user via the computing device or an additional computing device, The aforementioned method.

2. Using an automatic speech recognition (ASR) model, the audio data captured from the spoken utterance including the query is processed to generate ASR data for the query. Using a natural language understanding (NLU) model, the ASR data is processed to generate NLU data for the query. The method according to claim 1, further comprising:

3. Performing the first search on the first corpus data in order to identify one or more candidate responses to the query contained in the oral utterance is: The NLU data of the query and the indication of the corresponding vehicle sensor data instance are sent to the first corpus data, and the first search is performed on the first corpus data. Based on the content of the NLU data in response to the query, one or more candidate responses to the first search are identified. The method according to claim 2, including the method described in claim 2.

4. The first corpus data corresponds to user manual corpus data specific to the vehicle and provided by the vehicle's original equipment manufacturer (OEM), The method according to claim 1, wherein the second corpus data corresponds to additional corpus data that is not specific to the vehicle.

5. In response to the determination that the oral utterance containing the aforementioned query was not received within the threshold duration, The decision to perform the first search on the first corpus data and the second search on the second corpus data. The method according to claim 1, further comprising:

6. In response to the decision to perform the first search on the first corpus data and the second search on the second corpus data, Performing a first search on the first corpus data to identify one or more of the candidate responses to the query contained in the oral utterance, wherein the first search is based on (i) the query and (ii) the corresponding vehicle sensor data instance. To identify one or more of the candidate responses to the query contained in the oral utterance, a second search is performed on the second corpus data, wherein the second corpus data is added to the first corpus data, and the second search is performed such that (i) it is based on the query, but (ii) it is not based on the corresponding vehicle sensor data instance. The method according to claim 5, further comprising:

7. The method according to claim 6, further comprising selecting the given candidate response to be provided to the user via the computing device or an additional computing device, based on the ranking of the one or more candidate responses.

8. The ranking of the one or more candidate responses is as follows: One or more terms in the aforementioned query, The aforementioned corresponding vehicle sensor data instance, One or more user context signals characterizing the user's state in the vehicle, One or more vehicle context signals that characterize the state of the vehicle, or Application data associated with one or more applications accessible by the computing device or the additional computing device. The method according to claim 7, based on one or more of the above.

9. Deciding to perform the first search on the first corpus data but not the second search on the second corpus data is: The method according to claim 1, comprising determining that one or more terms of the query contained in the oral utterance are related to the corresponding vehicle sensor data instance.

10. The method according to claim 9, wherein, once the corresponding vehicle sensor data instance is generated, the computing device or the additional computing device provides an indication of the corresponding vehicle sensor data instance for presentation to the user.

11. Determining that one or more of the terms in the query contained in the oral utterance are related to the corresponding vehicle sensor data instance means The method according to claim 10, comprising determining that one or more terms of the query contained in the oral utterance target the indication of the corresponding vehicle sensor data instance provided for presentation to the user via the computing device or the additional computing device.

12. Based on processing the oral utterance including the query, further includes identifying the duration associated with the vehicle by the user, Performing speaker identification based on processing the aforementioned oral utterance and determining whether the user is a known user, In response to determining that the aforementioned user is a known user, Based on the user profile associated with the user, the duration to which the user was associated with the vehicle is identified. It further includes, If the duration is within the threshold duration, it is decided to perform the first search on the first corpus data, but not the second search on the second corpus data, or If the duration is greater than the threshold duration, the system includes deciding to perform the first search on the first corpus data and the second search on the second corpus data. The method according to claim 1.

13. Deciding to perform the first search on the first corpus data but not the second search on the second corpus data is: The determination of whether the duration associated with the user with the vehicle satisfies the threshold duration, In response to the user determining that the duration associated with the vehicle does not meet the threshold duration, The first search is performed on the first corpus data, but the second search is not performed on the second corpus data. The method according to claim 12, including the method described in claim 12.

14. In response to the user determining that the duration associated with the vehicle satisfies the threshold duration, (i) to perform the first search on the first corpus data to identify one or more of the candidate responses to the query contained in the oral utterance based on the query and (ii) the corresponding vehicle sensor data instance, (i) Performing the second search on the second corpus data to identify one or more of the candidate responses to the query contained in the oral utterance based on the query, wherein the second corpus data is added to the first corpus data to perform the second search. The method according to claim 13, further comprising:

15. The method according to claim 14, further comprising selecting the given candidate response to be provided to the user via the computing device or an additional computing device, based on the ranking of the one or more candidate responses.

16. A method implemented by one or more processors, Receiving an oral utterance from a user via a computing device, the oral utterance being provided while the user is inside the user's vehicle, The acquisition of a corresponding vehicle sensor data instance, wherein the corresponding vehicle sensor data instance is generated by one or more vehicle sensors of the user's vehicle, Processing the oral utterance to identify a plurality of candidate responses to the query contained in the oral utterance, Performing a first search on a first corpus data to identify one or more first candidate responses among the plurality of candidate responses to the query contained in the oral utterance, wherein the first search is based on (i) the query and (ii) the corresponding vehicle sensor data instance. To identify one or more second candidate responses among the plurality of candidate responses to the query contained in the oral utterance, a second search is performed on a second corpus data, wherein the second corpus data is added to the first corpus data, and the second search is performed such that (i) it is based on the query, but (ii) it is not based on the corresponding vehicle sensor data instance. Processing the aforementioned oral utterance, Ranking the aforementioned multiple candidate responses, Applying a bias to the ranking of the plurality of candidate responses towards one or more first candidate responses, Selecting a given candidate response from among the multiple candidate responses based on the biased ranking of the multiple candidate responses, To provide the given candidate response to the user via the computing device or an additional computing device. The method, including the method described above.

17. Using an automatic speech recognition (ASR) model, the audio data captured from the spoken utterance including the query is processed to generate ASR data for the query. Using a natural language understanding (NLU) model, the ASR data is processed to generate NLU data for the query. The method according to claim 16, further comprising:

18. Performing the first search on the first corpus data in order to identify one or more first candidate responses to the query contained in the oral utterance is: The NLU data of the query and the indication of the corresponding vehicle sensor data instance are sent to the first corpus data, and the first search is performed on the first corpus data. Based on the first content that responds to the NLU data of the query, one or more first candidate responses for the first search are identified. The method according to claim 17, including the method described in claim 17.

19. Performing the second search on the second corpus data in order to identify one or more second candidate responses to the query contained in the oral utterance is: The NLU data of the aforementioned query is submitted to the second corpus data, and the second search is performed on the second corpus data. Identifying one or more second candidate responses based on the second content responding to the NLU data of the query. The method according to claim 18, including the method described in claim 18.

20. The first corpus data corresponds to user manual corpus data specific to the vehicle and provided by the vehicle's original equipment manufacturer (OEM), The method according to claim 16, wherein the second corpus data corresponds to a web-based corpus data that is not specific to vehicles.

21. The ranking of the aforementioned multiple candidate responses is, One or more terms in the aforementioned query, The aforementioned corresponding vehicle sensor data instance, One or more user context signals characterizing the user's state in the vehicle, One or more vehicle context signals that characterize the state of the vehicle, or Application data associated with one or more applications accessible by the computing device or the additional computing device. The method according to claim 16, based on one or more of the above.

22. A method implemented by one or more processors, Receiving an oral utterance containing a query from a user via a computing device, wherein the oral utterance is received while the user is inside the user's vehicle, The acquisition of a corresponding vehicle sensor data instance, wherein the corresponding vehicle sensor data instance is generated by one or more vehicle sensors of the user's vehicle, Based on processing the oral utterance including the query, the duration for which the user was associated with the vehicle is identified, (i) the query, (ii) the corresponding vehicle sensor data instance, and / or (iii) whether the duration for which the user was associated with the vehicle does not meet a temporal threshold, to determine whether to perform at least the first search of a first search in a first corpus data and a second search in a second corpus data to identify one or more candidate responses to the query contained in the oral utterance, Identifying a corresponding timestamp associated with the corresponding vehicle sensor data instance, wherein the corresponding timestamp associated with the corresponding vehicle sensor data instance corresponds to the time the corresponding vehicle sensor data instance was generated. Based on the corresponding timestamp associated with the corresponding vehicle sensor data instance, it is determined whether the verbal utterance including the query was received within a threshold duration with respect to the time the corresponding vehicle sensor data instance was generated. In response to the determination that the oral utterance containing the query was received within the threshold duration, The decision is made to perform the first search on the first corpus data, but not to perform the second search on the second corpus data. In response to the decision to perform the first search on the first corpus data but not the second search on the second corpus data, (i) to perform the first search on the first corpus data to identify one or more candidate responses to the query contained in the oral utterance based on the query and (ii) the corresponding vehicle sensor data instance, To provide the user with a given candidate response from among the one or more candidate responses via the computing device or an additional computing device. Includes, Identifying the duration for which the user is associated with the vehicle means that Performing speaker identification based on processing the aforementioned oral utterance and determining whether the user is a known user, In response to determining that the aforementioned user is a known user, This includes identifying the duration to which the user was associated with the vehicle, based on the user profile associated with the user. The aforementioned method.

23. A method implemented by one or more processors, Receiving an oral utterance from a user via a computing device, the oral utterance being provided while the user is inside the user's vehicle, The acquisition of a corresponding vehicle sensor data instance, wherein the corresponding vehicle sensor data instance is generated by one or more vehicle sensors of the user's vehicle, Identifying a corresponding timestamp associated with the corresponding vehicle sensor data instance, wherein the corresponding timestamp associated with the corresponding vehicle sensor data instance corresponds to the time the corresponding vehicle sensor data instance was generated. (i) determining whether to perform at least the first search of a first search in a first corpus data and a second search in a second corpus data to identify one or more candidate responses to the query contained in the oral utterance, based on (i) the query, (ii) the corresponding vehicle sensor data instance, and / or (iii) the corresponding timestamp associated with the corresponding vehicle sensor data instance, Based on the corresponding timestamp associated with the corresponding vehicle sensor data instance, it is determined whether the verbal utterance including the query was received within a threshold duration with respect to the time the corresponding vehicle sensor data instance was generated. In response to the determination that the oral utterance containing the query was received within the threshold duration, The decision is made to perform the first search on the first corpus data, but not to perform the second search on the second corpus data. In response to the decision to perform the first search on the first corpus data but not the second search on the second corpus data, (i) to perform the first search on the first corpus data to identify one or more candidate responses to the query contained in the oral utterance based on the query and / or (ii) the corresponding vehicle sensor data instance, To provide the user with a given candidate response from among the one or more candidate responses via the computing device or an additional computing device. The method, including the method described above.

24. One or more processors, When executed, the memory stores instructions that cause one or more processors to perform the operation described in any one of claims 1 to 23. A system that includes this.

25. A non-temporary computer-readable storage medium that stores instructions, when executed, causing one or more processors to perform the operations described in any one of claims 1 to 23.

26. One or more processors, When executed, the memory stores instructions that cause one or more processors to perform the operation described in any one of claims 1 to 23. In-vehicle computing devices, including

Citation Information

Patent Citations

  • Information retrieval apparatus, information storage device and program

    JP2011210136A

  • Acquiring response information from multiple corpora

    JP2020526812A

  • Vehicle personal assistant

    US20140136013A1

  • Intelligent User Manual System for Vehicles

    US20210023945A1