Selective interaction of a robotic device with an additional computing device
The robotic computing device optimizes task delegation among nearby devices based on capability and efficiency, addressing inefficiencies in static assistants by navigating for additional information or task completion, enhancing resource utilization and accuracy.
Patent Information
- Application Number
- JP2024534079
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-07
- Filing Date
- 2022-10-07
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2042-10-07
AI Technical Summary
Existing automated assistants are often static and inefficient in fulfilling user requests, especially when users move around, and may delegate tasks to less suitable devices due to software provider or location constraints, leading to resource wastage.
A robotic computing device that can interact with nearby devices to delegate tasks based on capability and efficiency, navigating to other devices for additional information or task completion, using techniques similar to human-computer interaction.
Enhances task completion efficiency by utilizing suitable devices, conserving resources and minimizing time, while ensuring accurate and timely fulfillment of user requests.
Smart Images

Figure 0007733829000001 
Figure 0007733829000002 
Figure 0007733829000003
Abstract
Description
[Background technology]
[0001] Automated assistants have been widely deployed in many homes, providing hands-free access to information from the Internet and controlling other peripherals inside and outside the home. Often, assistant-enabled devices are static and may not have the ability to move to different locations within the home. As a result, operations performed by an automated assistant may be geographically restricted depending on where the user places the assistant-enabled device during installation. This may not be particularly inefficient when performing certain tasks, such as when a user issues an internet search query (e.g., "Who is Joseph Fourier?"). However, other tasks involving communication between users and / or rendering output to a user while on the move may prove inefficient or ineffective when performed by a static device or network of static devices. For example, a user listening to a phone call rendered by a kitchen assistant device may lose the audio if they move from the kitchen to a hallway where no assistant-enabled device is present. As a result, the user may have to ask the caller to repeat what they said and / or pause the call when they need to leave the room, wasting time and computational resources.
[0002] In some cases, an assistant-enabled device may be able to respond to a query from a user by having another assistant-enabled device render output responsive to the query. However, delegating the performance of certain operations in this manner may be ineffective in a multi-assistance environment where multiple assistant devices are associated with different software providers and / or user accounts. As a result, the performance of certain operations may be delegated to a less suitable assistant device, even though other more suitable devices (e.g., devices with better sound quality, a stronger signal, more efficient power usage, etc.) are available. For example, communication tasks between users may be performed by default by the assistant device that receives the corresponding request from the user. However, if other more suitable devices are available to fulfill such requests and / or if the default device is not positioned to effectively communicate with any of the specific users, the lack of device capacity to interact may result in further wasted resources. Summary of the Invention [Means for solving the problem]
[0003] Embodiments described herein relate to a robotic computing device that can interact with other nearby devices (possibly using techniques common to human users) to facilitate fulfilling requests made by a user to the robotic computing device. The robotic computing device can, for example, render voice commands to nearby Assistant-enabled devices to fulfill requests made by a user to the robotic computing device. For example, a user may make a voice utterance to the robotic computing device such as, "Can you help Emma find her toy car?", which may be a request for the robotic computing device to identify a particular object in the user's home. In response to receiving the voice utterance, the robotic computing device may determine that a search for information stored on the robotic computing device to identify an appropriate description of "toy car" does not result in an accurate description. For example, the robotic computing device has internet search capabilities, but search results for the phrases "Emma's toy car" or "toy car" may not provide an accurate description of the object the user is referring to. Thus, the robotic computing device, with prior or current permission from the user, may in turn generate output commands to be provided to one or more other computing devices within the user's home to obtain more precise details about the object the user is referring to.
[0004] In some embodiments, the robot computing device may generate natural language content, such as "Assistant, what is 'Emma's toy car'?" and then render the natural language content as an audio output to a nearby computing device. In response, the nearby computing device (e.g., another Assistant-enabled device) may render an image of "Emma's toy car" on a display panel. One or more camera images of the displayed image may be captured by a camera on the robot computing device with prior permission from the user, and the robot computing device may use these camera images to estimate a likely location of "Emma's toy car." For example, prior images captured by the robot computing device and associated with locations on the home graph may be processed to determine whether an object in a camera image is one captured in any of the prior images. If a particular prior image is determined to include the object (e.g., "Emma's toy car"), the robot computing device may identify and use a saved map location associated with the particular prior image. The robot computing device may then navigate from the robot computing device's current location (e.g., the location from which the robot computing device emitted the audible output to a nearby device) to a map location corresponding to the object in the particular camera image. Alternatively or additionally, the robotic computing device may generate additional natural language content representing the location on the map and render another output representing the additional natural language content to the user (e.g., an audible or visual output that says, "The toy car is under the kitchen table").
[0005] In some embodiments, a robotic computing device may determine whether to delegate an action to another nearby device and may perform such delegation using one or more techniques similar to human-computer interaction. That is, if the robotic computing device determines that another computing device may be better suited to perform a requested action, the robotic computing device may issue a human-perceivable command to the other computing device. For example, a user may make a vocal utterance to the robotic computing device, such as, "Play some music while I cook." In response, the robotic computing device may determine that a plugged-in device is better suited to fulfill this request, at least in part because the robotic computing device is operating using battery power. Based on this determination, the robotic computing device may identify a nearby device that can more efficiently fulfill the user's request.
[0006] For example, the robot computing device may identify kitchen smart displays (i.e., standalone display devices) within a threshold distance from the user to render audio content. In some embodiments, the robot computing device may identify the device name or type of kitchen smart display and render commands for the robot computing device based on the name or type of the kitchen smart display. For example, based on identifying the type of device nearby, the robot computing device may identify a call phrase for that particular type of device (e.g., “OK, smart device…”). Using this call phrase and with prior permission from the user, the robot computing device may render an audible output such as “OK, smart device, play some music while I cook.” In response, the kitchen smart display may begin rendering music. In this way, the robot computing device can act as an interface between the user and all of its smart devices, even if the kitchen smart display is not specifically pre-configured to receive task delegation from the robot computing device. When task delegation is performed based on capability and / or efficiency, smart devices in the user's home can follow the instructions of the robot computing device to conserve resources such as bandwidth and power.
[0007] In some embodiments, a robotic computing device may leverage the capabilities of another device if it determines that another device can complete a requested task faster than the robotic computing device itself. This may occur when a standalone device and / or other robotic computing devices are determined to be closer to the task location than the robotic computing device that originally received the request to complete the task. For example, a user ("Kaiser") may speak an utterance to a robotic computing device, such as, "Tell Karma to turn off the heater upstairs." The robotic computing device may be located on the first floor of the user's home along with the standalone computing device, and another standalone computing device may be located on the second floor of the home closer to another user ("Karma"). Upon receiving the voice utterance from the user, the robotic computing device may determine, with prior permission from one or more people in the home, that the other users are on a different floor of the home from the robotic computing device. Based on this determination, the robotic computing device may cause another standalone computing device on the second floor to send a message to the other users.
[0008] In some embodiments, a robotic computing device can cause another standalone computing device to send a message in a variety of different ways. For example, a robotic computing device may communicate with a standalone computing device on the first floor to cause it to send a message (e.g., "Kaiser wants you to turn off the heater") to another standalone computing device on the second floor. Communication between the robotic computing device and the standalone computing device can occur via a local area network (LAN), a wide area network (WAN) such as the Internet, audible or inaudible frequencies, Bluetooth communication, and / or any other medium used for communication between devices. For example, a robotic computing device may generate natural language content corresponding to a command to be sent to a standalone speaker device on the first floor. The natural language content may be embodied in an audible or inaudible message to the standalone speaker device (e.g., "OK, smart device, please send a message to Karma telling her to turn off the heater on the second floor."). In response, another standalone computing device on the second floor may render an audible output and / or a visual message (e.g., "Please turn off the heater on the second floor"). Providing a robotic computing device that can interface with other computers in a user's home can allow for more efficient utilization of resources within the home while minimizing the time required to complete requested tasks.
[0009] The above description is provided as a summary of some embodiments of the present disclosure. These and other embodiments are described in more detail below.
[0010] Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., central processing unit (CPU)(s), graphics processing unit (GPU)(s), and / or tensor processing unit (TPU)(s)) to perform methods such as one or more of the methods described above and / or elsewhere herein. Still other embodiments may include one or more computer systems including one or more processors operable to execute the stored instructions to perform methods such as one or more of the methods described above and / or elsewhere herein.
[0011] As will be apparent, all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein, for example, all combinations of subject matter recited in the claims at the end of this disclosure are considered to be part of the subject matter disclosed herein. [Brief explanation of the drawings]
[0012] [Figure 1A] FIG. 1 illustrates a robotic computing device that operates to assist a user with the assistance of other nearby computing devices. [Figure 1B] FIG. 1 illustrates a robotic computing device that operates to assist a user with the assistance of other nearby computing devices. [Figure 1C] FIG. 1 illustrates a robotic computing device that operates to assist a user with the assistance of other nearby computing devices. [Figure 1D] FIG. 1 illustrates a robotic computing device that operates to assist a user with the assistance of other nearby computing devices. [Figure 1E]FIG. 1 illustrates a robotic computing device that operates to assist a user with the assistance of other nearby computing devices. [Figure 2] FIG. 1 illustrates a system for operating a robotic computing device that can seek additional information from other nearby devices in order to fulfill requests and / or delegate certain tasks to the other nearby devices. [Figure 3] FIG. 1 illustrates a method for controlling a robotic computing device to carry out a request from a user by delegating performance to one or more other computing devices and / or by gathering additional information from one or more other computing devices. [Figure 4] FIG. 1 is a block diagram of an example computer system. DETAILED DESCRIPTION OF THE INVENTION
[0013] 1A, 1B, 1C, 1D, and 1E illustrate views 100, 120, 140, 160, and 180 of a robotic computing device 104 operating to assist a user 102 with the assistance of other nearby computing devices. The robotic computing device 104 may determine whether to seek assistance from additional computing devices based on various factors, such as the confidence that the robotic computing device 104 can accurately fulfill a request from the user, the amount of energy expected to be consumed by one or more devices, the time it takes to fulfill a request by one or more devices, the amount of processing bandwidth consumed by one or more devices, and / or other factors that may be considered when delegating a task to and / or being delegated by a particular device. For example, the robotic computing device 104 and the user 102 may be located in an environment 106 (e.g., the user's 102's home) when the user 102 provides a voice utterance 110 (e.g., "Can you help me find Luke's diary?"). The robotic computing device 104 can provide a response output 108 in response to the vocal utterance 110 (e.g., "Yes, I can help you find Luke's diary").
[0014] In some embodiments, in response to receiving the vocal utterance 110, the robot computing device 104 may process audio data corresponding to the vocal utterance 110 to identify one or more operations to perform in response to the vocal utterance 110. For example, the robot computing device 104 may determine that the user 102 requests assistance in locating an object within the environment 106. In some embodiments, this request may be fulfilled by at least executing an intent having a slot value associated with the object to be located. For example, the intent's slot value may include an image of the object. However, even though the robot computing device 104 has access to multiple images that may resemble and / or be associated with the object, the robot computing device 104 may determine that the multiple images do not meet a threshold confidence level. As a result, the robot computing device 104 may determine whether to delegate fulfillment of the request to another device and / or whether to seek additional information from another device and / or user within (or outside) the environment 106.
[0015] For example, in response to determining that the robot computing device 104 may be unable to fulfill the request with a threshold confidence, the robot computing device 104 may determine whether another device is available to assist in fulfilling the request. For example, the robot computing device 104 may use sensor data and / or other data available to the robot computing device 104 to determine whether the environment 106 includes one or more other devices that can be invoked to provide additional information. In some embodiments, such data may include image data and / or audio data captured by the robot computing device 104 with prior permission from a user in the home. For example, images captured by the robot computing device 104 may be processed to determine whether the environment 106 includes one or more other computing devices and / or to determine the type of one or more other computing devices in the environment 106. When the robot computing device 104 identifies a smart display device 124 and / or a smart speaker device 122 (“smart” may indicate that the device has the capability to access the Internet and respond to user input), the robot computing device 104 may determine whether an invoke phrase is required to invoke the identified device. For example, the robotic computing device 104 may perform an internet search and / or other database search to determine how to request additional information from a particular device.
[0016] As shown in view 120 of FIG. 1B , the robotic computing device 104 may navigate to an area of the environment 106 where the smart display device 124 resides and may request additional information from the smart display device 124 using the determined call phrase. For example, the robotic computing device 104 may use its wheel(s) and / or track(s) to navigate to an area of the environment, possibly using a pre-created environment map and / or navigation method(s) (e.g., simultaneous map and localization (SLAM)). While FIG. 1B and other figures illustrate a particular robotic computing device 104, it should be understood that embodiments disclosed herein may be implemented in additional or alternative mobile robotic computing devices. For example, they may be implemented in other mobile computing device(s) capable of self-navigating using wheel(s), track(s), leg(s) (e.g., a multi-legged robot), rotors (e.g., an unmanned aerial vehicle robot), and / or wing(s). Additionally, the robot computing device 104 may generate content for the robot input provided by the robot computing device 104 to the smart display device 124. The content may be generated using one or more heuristic processes and / or one or more trained machine learning models (e.g., language models). For example, the robot computing device 104 may not have a threshold level of confidence regarding the appearance of a particular object, so the robot computing device 104 may generate content that embodies a request to describe the characteristics of the particular object. Following the example above, the robot computing device 104 may render audio output 130 including content such as "Assistant, ...", which may mean an invocation phrase, or content such as "...what does Luke's diary look like?", which may mean a command requesting additional information from the smart display device 124.
[0017] In response to the audio output 130, the smart display device 124 may render search results 142 on a display interface 144 of the smart display device 124, as shown in view 140 of FIG. 1C . Alternatively or additionally, the smart display device 124 may provide an audio response 146, such as, "Here are the search results for 'Luke's Diary'," followed by an audio description of the characteristics of the object for which the robot computing device 104 is querying. In some embodiments, as the smart display device 124 is rendering the search results 142 and / or the audio description of the object, the robot computing device 104 may capture data based on the output that the smart display device 124 is rendering. For example, the robot computing device 104 may use one or more cameras 166 and / or one or more microphones to capture the output that the smart display device 124 is rendering. The smart display device 124 may capture an image of the search results 142. The image may include images of diaries recently searched for on the internet by one or more users. This image may be used by the robotic computing device 104 to fulfill a request from the user 102 .
[0018] In some embodiments, data captured by the robotic computing device 104 may be compared with private and / or public home knowledge graph data, with prior permission from the user(s), to determine the location of an object queried by the user 102. Values stored in the knowledge graph may include text values (e.g., names of objects, names of places, other textual descriptors of entities), numerical values (e.g., entity types, usage data, age, height, weight, other characteristic data, other numerical information associated with entities), or pointers to user-specific values (e.g., memory locations for entities in the user's knowledge graph, memory locations relating two or more entities in the user's knowledge graph, etc.). That is, values specific to a user and / or environment (e.g., a particular home) may take various forms and may be specific to fields of a personal record defined by a record schema. The value may represent actual information specific to the user, or may be a designation of a memory location and / or device from which information specific to the user and / or environment can be obtained.
[0019] In some cases, a comparison of the data presented by the smart display device 124 with the data graphed in the personal knowledge graph and / or home knowledge graph may identify information that may assist the robot computing device 104 in fulfilling a request from the user 102. For example, the robot computing device 104 may determine a location of the identified object based on similar objects that have been captured and location data stored in the home knowledge graph of the user 102. Once the location is identified, the robot computing device 104 may optionally determine whether the location was determined with a threshold certainty or confidence. If the robot computing device 104 determines with a threshold confidence that the object identified by the user (e.g., Luke's diary) is located at that location, the robot computing device 104 may render instructions to the user 102.
[0020] For example, as shown in view 160 of FIG. 1D , the robot computing device 104 may navigate from an area in the environment 106 where the user 102 requested help from the robot computing device 104 to another area. The other area may be a location predicted by the robot computing device 104 with prior permission from the user 102 and to which the user 102 has traveled. Depending on the output rendered to the user 102, the robot computing device 104 may navigate to the user 102's location and / or request another computing device to render the output to the user 102. For example, if the user 102 is predicted to be near another computing device, the robot computing device 104 may communicate with the other computing device (e.g., via a Wi-Fi network) and have the other computing device render the output to the user 102. Otherwise, the robot computing device 104 may navigate to a location near the user 102 and render an audible output 162 and / or a visual output. For example, once the robotic computing device 104 determines the information to fulfill the request from the user 102, the robotic computing device 104 may render an audible output 1062 such as, "Luke's diary is in the kitchen. Let me take you to Luke's diary." In response, the user 102 can provide an affirmative answer 164 such as, "Okay, thanks," confirming that the user 102 wishes to be guided to the location of the specified object.
[0021] In response to the user 102 confirming the robot computing device 104's offer to guide them to the object, the robot computing device 104 may move to another area within the environment 106. This other area may be, for example, a kitchen that includes a counter on which the desired object is located. In some cases, as shown in view 180 of FIG. 1E , the user 102 may issue another request to the robot computing device 104. The other request may be, for example, a vocal utterance 182 such as, "Tell Luke that his diary is in the kitchen." This request may be the user 102 asking the robot computing device 104 to convey a message to another user. In some embodiments, the robot computing device 104 may process this input from the user 102 and determine whether to delegate fulfillment of the request. For example, using home knowledge graph data and / or other data, the robot computing device 104 may determine whether one or more other devices can more effectively fulfill the request from the user 102.
[0022] In some cases, the robot computing device 104, with prior permission from the user(s), may determine a predicted location of the user Luke and estimate whether the robot computing device 104 can effectively communicate with Luke or whether another device can more effectively communicate with Luke. For example, the robot computing device 104 may determine, based on personal knowledge graph data, that the user Luke is playing with his cell phone in his room and that a smart speaker device is in his room. Based on this determination, the robot computing device 104 may determine that without the assistance of another device, it would consume more energy and / or take longer for the robot computing device 104 to communicate with Luke. Thus, the robot computing device 104 may determine that a smart speaker device 188 is nearby and, based on the type of smart device, may invoke the smart speaker device 188 to send a message to the user Luke. For example, the robot computing device 104 may generate audible output content such as, "Assistant, tell Luke that his diary is in the kitchen." This audible output may enable the smart speaker device 188 and / or one or more other computing devices to render output 190 such as, "Luke, your diary is in the kitchen."
[0023] In some embodiments, the robot computing device 104 may use the smart speaker device 188, the smart display device 124, and / or other computing devices to confirm that the object 186 identified by the robot computing device 104 (e.g., Luke's diary) is the object that the user 102 was trying to find. For example, the robot computing device 104 may identify the object 186, manipulate an arm or other part of the robot computing device 104 to pick up the object 186, and bring the object 186 to a location on the smart display device 124 or the smart speaker device 188 (e.g., if the smart speaker device 188 includes a camera). The robot computing device 104 may then request the smart display device 124 (or other device) to confirm whether the object 186 is the object that the user 102 was referring to.
[0024] For example, the robotic computing device 104 may generate a command including an invocation phrase and content requesting the smart display device 124 to capture an image of what the robotic computing device 104 is holding. For example, the command may be, "Smart device, is this Luke's diary?" In some embodiments, the command may be generated based on a prior interaction between the user 102 and the smart display device 124. Such a prior interaction informs the robotic computing device 104 of the capabilities of the smart display device 124. Alternatively or additionally, the command may be generated based on a search of a database, the internet, or other information source to determine the type of command to which the smart display device 124 will respond.
[0025] In some embodiments, the robot computing device 104 may confirm to the user 102 or another user that the object 186 identified by the robot computing device 104 is the object the user 102 refers to. For example, the robot computing device 104 may initiate a request to have another user (e.g., Luke) confirm the name of the object 186 by providing a command to the smart display device 124 to conduct a video call with the robot computing device 104. For example, the robot computing device 124 may generate a command such as "Smart device, call Luke on a video call" and render the command as an audio output to the smart display device 124. In response, the smart display device 124 may initiate a video call between the robot computing device 124 and the other user (with prior permission from the other user). In some cases, the other user may not be present in the environment 106, but the robot computing device 104 may still initiate communication with the other user via another computing device in the environment 106. When a video call is initiated on the smart display device 124, the robotic computing device 104 may hold up the object 186 and request the other user to confirm the name of the object 186 (e.g., "Hi, Luke. Is this your diary?"). If the other user responds by confirming the name of the object 186, the robotic computing device 104 may consider the request from the user 102 to be fulfilled.
[0026] FIG. 2 illustrates a system 200 for operating a robotic computing device (e.g., computing device 202) that can seek additional information from other nearby devices to fulfill requests and / or delegate certain tasks to the other nearby devices. The robotic computing device may include an automated assistant 204. The automated assistant 204 may operate as part of an assistant application provided on one or more computing devices, such as computing device 202 and / or a server device. A user can interact with the automated assistant 204 through assistant interface(s) 220. The assistant interface(s) 220 can be a microphone, a camera, a touchscreen display, a user interface, and / or other device capable of providing an interface between a user and an application. For example, a user can initialize the automated assistant 204 and cause the automated assistant 204 to initiate one or more actions (e.g., provide data, control a peripheral, access an agent, generate input and / or output, etc.) by providing verbal, textual, and / or graphical input to the assistant interface 220. Alternatively, automated assistant 204 may be initialized based on processing context data 236 using one or more trained machine learning models. Context data 236 may characterize one or more characteristics of the environment in which automated assistant 204 is accessible and / or one or more characteristics of a user who is predicted to intend to interact with automated assistant 204. Computing device 202 may include a display device. The display device may be a display panel including a touch interface for receiving touch inputs and / or gestures that enable a user to control application 234 of computing device 202 via the touch interface.In some embodiments, computing device 202 may not have a display device and therefore may provide audible user interface output without providing graphical user interface output. Additionally, computing device 202 may provide a user interface such as a microphone for receiving spoken natural language input from a user. In some embodiments, computing device 202 may include a touch interface and may lack a camera, but may possibly include one or more other sensors.
[0027] Computing device 202 and / or other third-party client devices may communicate with a server device over a network, such as the Internet. Additionally, computing device 202 and other computing devices may communicate with each other over a local area network (LAN), such as a Wi-Fi network. Computing device 202 may offload computational tasks to a server device to conserve computing resources of computing device 202. For example, the server device may host automated assistant 204 and / or computing device 202 may send input received at one or more assistant interfaces 220 to the server device. However, in some embodiments, automated assistant 204 may be hosted on computing device 202, and various processes that may be associated with automated assistant operation may be executed on computing device 202.
[0028] In various embodiments, all or less than all of the automated assistant 204 may execute on the computing device 202. In some of these embodiments, some portions of the automated assistant 204 may execute via the computing device 202 and may interface with a server device that can execute other portions of the automated assistant 204. The server device may, in some cases, serve multiple users and their associated assistant applications via multiple threads. In embodiments in which all or less than all of the automated assistant 204 executes via the computing device 202, the automated assistant 204 may be an application separate from the operating system of the computing device 202 (e.g., installed “on top of” the operating system) or may be executed directly by the operating system of the computing device 202 (e.g., an operating system application but considered integral to the operating system).
[0029] In some embodiments, the automated assistant 204 may include an input processing engine 206. The input processing engine 206 may employ multiple different modules to process input and / or output for the computing device 202 and / or the server device. For example, the input processing engine 206 may include a speech processing engine 208, which may process speech data received at the assistant interface 220 to identify text contained in the speech data. The speech data may be transmitted from the computing device 202 to the server device, for example, to conserve computational resources of the computing device 202. Additionally or alternatively, the speech data may be processed exclusively on the computing device 202.
[0030] The process of converting the speech data to text may include speech recognition algorithms, which may employ neural networks, and / or statistical models to identify groups of speech data that correspond to words or phrases. The text converted from the speech data may be parsed by data parsing engine 210 and made available to automated assistant 204 as text data. Such text data may be used to generate and / or identify command phrase(s), intent(s), action(s), slot value(s), and / or other user-specified content. In some embodiments, output data provided by data parsing engine 210 may be provided to parameter engine 212 to determine whether the user has provided input corresponding to a particular intent, action, and / or routine executable by automated assistant 204 and / or by an application or agent accessible via automated assistant 204. For example, assistant data 238 may be stored on server device and / or computing device 202 and may include data defining one or more actions executable by automated assistant 204 and parameters required to perform the actions. The parameter engine 212 may generate one or more parameters for the intent, action, and / or slot value and may provide the one or more parameters to the output generation engine 214. The output generation engine 214 may use the one or more parameters to communicate with the assistant interface 220 to provide output to the user and / or to communicate with one or more applications 234 to provide output to the one or more applications 234.
[0031] In some embodiments, the automated assistant 204 may be an application that can be installed “on top of” and / or that itself forms part of (or entirely forms) the operating system of the computing device 202. The automated assistant application may include and / or have access to on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition may be performed using an on-device speech recognition module that processes speech data (detected by the microphone(s)) using an end-to-end speech recognition machine learning model stored locally on the computing device 202. The on-device speech recognition generates recognized text for speech utterances (if any) present in the speech data. Also, for example, on-device natural language understanding (NLU) may be performed using an on-device NLU module. The on-device NLU module processes recognized text generated using on-device speech recognition and, optionally, processes contextual data to generate NLU data.
[0032] The NLU data may include the intent(s) corresponding to the voice utterance and, possibly, parameter(s) of the intent(s) (e.g., slot values). On-device fulfillment may be performed using an on-device fulfillment module that uses the NLU data (from the on-device NLU) and, possibly, other local data, to determine the action(s) to perform to resolve the intent(s) of the voice utterance (and, possibly, the parameter(s) of the intent). This may include determining local and / or remote responses (e.g., answers) to the voice utterance, determining interaction(s) with locally installed application(s) to perform based on the voice utterance, determining command(s) to send (directly or via corresponding remote system(s)) to Internet of Things (IoT) device(s) based on the voice utterance, and / or determining other resolution action(s) to perform based on the voice utterance. On-device fulfillment may then initiate local and / or remote execution / execution of the determined action(s) to resolve the voice utterance.
[0033] In various embodiments, remote speech processing, remote NLU, and / or remote fulfillment may be at least selectively utilized. For example, recognized text may at least selectively be sent to a component(s) of a remote automated assistant for remote NLU and / or remote fulfillment. For example, recognized text may possibly be sent for remote fulfillment in parallel with on-device fulfillment or in response to a failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized due to at least reduced latency incurred when resolving a speech utterance (because client-server round-trip(s) are not required to resolve the speech utterance). Furthermore, on-device functionality may be the only functionality available in situations where there is no network connectivity or where network connectivity is limited.
[0034] In some embodiments, computing device 202 may include one or more applications 234. The one or more applications 234 may be provided by a third-party entity different from the entity that provided computing device 202 and / or automated assistant 204. An application state engine of automated assistant 204 and / or computing device 202 may access application data 230 to determine one or more actions performable by one or more applications 234 and may further determine the state of each of the one or more applications 234 and / or the state of each device associated with computing device 202. A device state engine of automated assistant 204 and / or computing device 202 may access device data 232 to determine one or more actions performable by computing device 202 and / or one or more devices associated with computing device 202. Furthermore, application data 230 and / or other data (e.g., device data 232) may be accessed by automated assistant 204 to generate context data 236. The context data 236 may characterize the context in which a particular application 234 and / or device is running and / or the context in which a particular user is accessing the computing device 202 (accessing the application 234 and / or other device or module).
[0035] While one or more applications 234 are executing on the computing device 202, the device data 232 may characterize the current operational state of each application 234 executing on the computing device 202. Additionally, the application data 230 may characterize one or more features of the executing applications 234, such as the content of one or more graphical user interfaces rendered at the direction of the one or more applications 234. Alternatively or additionally, the application data 230 may characterize action schemas that can be updated by the respective applications and / or automated assistant 204 based on the respective applications' current operational state. Alternatively or additionally, one or more action schemas of one or more applications 234 may remain static but may be accessed by the application state engine to determine appropriate actions to initiate via the automated assistant 204.
[0036] Computing device 202 may further include assistant invocation engine 222. Assistant invocation engine 222 may use one or more trained machine learning models to process application data 230, device data 232, context data 236, and / or other data accessible to computing device 202. Assistant invocation engine 222 may process this data to determine whether to wait for the user to explicitly speak an invocation phrase to invoke automated assistant 204 or whether to consider the data to indicate the user's intent to invoke the automated assistant instead of having the user explicitly speak an invocation phrase. For example, one or more trained machine learning models may be trained using instances of training data. The instances of training data are based on scenarios in which a user is in an environment with multiple devices and / or applications exhibiting various operating states. The instances of training data may be generated to obtain training data that characterizes contexts in which a user invokes an automated assistant and other contexts in which the user does not invoke an automated assistant. Once one or more trained machine learning models have been trained based on these training data instances, the assistant invocation engine 222 may cause the automated assistant 204 to detect or limit the detection of voice invocation phrases from the user based on contextual and / or environmental features.
[0037] In some embodiments, assistant invocation engine 222 may process data to determine how to invoke a nearby device that provides access to an instance of automated assistant 204 and / or another assistant application. For example, image data captured by a camera of computing device 202 may be processed to determine the type of computing device located in the same environment as computing device 202. Computing device 202 may then determine that this type of computing device provides access to a particular automated assistant that can be invoked using a particular invocation phrase. Based on this determination, computing device 202 may initialize output generation engine 214 to render an invocation phrase when computing device 202 determines to seek additional information from a nearby device and / or delegate a task to a nearby device. Output generation engine 214 may then render output (e.g., an audible output, a Bluetooth command, a non-audible request, etc.) that includes the invocation phrase and content of the request generated by computing device 202.
[0038] In some embodiments, the system 200 may include a delegation engine 216 that can determine whether to delegate the fulfillment of a request and / or a portion of a request to another computing device. In some embodiments, the delegation engine 216 may process data to determine the state of devices and / or applications associated with a user to which a particular task can be delegated, with prior permission from the user. The data processed by the delegation engine 216 may be selected according to the request provided by the user, allowing the computing device 202 to customize the delegation of a particular task for each request. For example, when the delegation engine 216 is determining whether to delegate a task rendering music for a user, the data may indicate whether a particular interface of the device is being used. Alternatively or additionally, the delegation engine 216 may process power usage data to identify nearby devices that are connected to utility power or using battery power when determining whether to delegate various types of requests. In this manner, the computing device 202 can avoid delegating certain tasks to nearby devices that may not have the optimal power source to complete the task.
[0039] In some embodiments, the delegation engine 216 may access personal knowledge graph data and / or home graph data of one or more different users (with prior permission from the users) to make a decision about delegating a particular task to fulfill a request. For example, the home graph data may indicate, with prior permission from the user, the location of particular objects and / or features within the user's home, the status of particular devices (e.g., processing bandwidth, signal strength, battery charge), and / or other properties of devices and objects. Alternatively or additionally, the home graph data may indicate the status of particular devices and / or applications being accessed within the home. The personal knowledge graph data may also be used by the delegation engine 216 to determine whether to delegate a particular task to a particular device. For example, the personal knowledge graph data may indicate the user's preferences for having particular devices perform particular operations and / or other preferences that can be stored on the device. The computing device 202 may use such data to delegate a task that the user prefers to be performed on another device (even though the user has issued a corresponding request to the computing device 202).
[0040] In some embodiments, data generated by the robot computing device may be used to update the home graph data and / or personal knowledge graph data. For example, image data and / or audio data captured by the robot computing device with prior permission from the user as the robot computing device moves through the environment may be used to update the home graph data and / or personal knowledge graph data. In some embodiments, the robot computing device may proactively generate such data if a particular portion of the home graph data and / or personal knowledge graph data has not been updated for at least a threshold period of time. Alternatively or additionally, the robot computing device may update data regarding a particular object and / or a particular portion of the environment if the robot computing device has not observed the particular object and / or a particular portion of the environment for at least a threshold period of time.
[0041] In some embodiments, the computing device 202 may include an information request engine 218. The information request engine 218 may process data to determine whether to require additional information to fulfill a request from a user. If the information request engine 218 determines that a particular request cannot be fulfilled with a threshold level of confidence without requesting at least the additional information, the additional information may be requested from another device and / or application. In some embodiments, the decision to request additional information and / or to delegate a particular task may be based on one or more heuristic processes and / or one or more trained machine learning models. For example, when a user submits a request to the computing device 202, the delegation engine 216 and / or the information request engine 218 may generate a confidence metric that indicates whether the task needs to be delegated and / or whether additional information needs to be sought to fulfill the request.
[0042] For example, inputs embodying the request may be processed to determine reliability metrics for fulfilling the request. Alternatively or additionally, candidate responses may be generated by the computing device 202 and processed to determine reliability metrics for the candidate responses. If any of the reliability metrics do not satisfy the reliability metrics for delegating the task, the delegation engine 216 may initiate a process to determine an appropriate device to delegate the task to. If any of the reliability metrics do not satisfy the reliability metrics for requesting additional information, the information request engine 218 may initiate a process to determine an appropriate device from which to ask for the additional information. In some embodiments, the information request engine 218 may first determine whether additional information needs to be requested for a particular request, and then the delegation engine 216 may determine whether one or more tasks for the request need to be delegated to another device. For example, the computing device 202 may receive a request to send a message to another device, and the information request engine 218 may initially determine, with a threshold confidence, that no additional information needs to be requested to fulfill the request. Subsequently, or simultaneously, the delegation engine 216 may determine that a particular computing device is more suitable to perform the task of sending the message. The delegation engine 216 may then generate data used by the assistant invocation engine 222 to output requests from the computing device 202 to a particular computing device to cause the particular computing device to perform the request (e.g., send a message).
[0043] FIG. 3 illustrates a method 300 for controlling a robotic computing device to fulfill a request from a user by delegating performance to one or more other computing devices and / or by gathering additional information from one or more other computing devices. Method 300 may be performed by one or more computing devices, applications, and / or any other device or module that may be associated with an automated assistant. Method 300 may include an operation 302 for determining whether the robotic computing device receives user input. The user input may, for example, be a request to the robotic computing device to provide the location of an object in the user's home. For example, the user may utter a voice utterance such as, "Where is Maggie's drone?", which refers to a toy drone owned by another user who is in the same environment (e.g., home) as the user. The user provides the voice utterance to the robotic computing device to actively recruit the robotic computing device to assist the user and other users in finding the object (e.g., "Maggie's drone").
[0044] Method 300 may proceed from operation 302 to operation 304. Operation 304 may include determining whether the robot computing device needs to delegate fulfillment of the request to another device in the environment. In some cases, the environment may be a user's home and include multiple different devices that provide access to one or more different automated assistant applications. For example, the environment may include an office with a first assistant device and a kitchen with a second assistant device. When the robot computing device receives user input, the robot computing device may process the user input and determine that the user has requested that the robot computing device locate an object. Based on this determination, the robot computing device may estimate the location of the object and a magnitude of a confidence metric for the location estimate.
[0045] In some embodiments, location estimation may be based on an image search performed in response to receiving user input from a user. One or more image results identified from the image search may be compared, with the user's prior permission, to one or more images captured by the robotic computing device as the robotic computing device moves through the environment. The robotic computing device and / or another device (e.g., a server device) may determine whether an object identified by the user corresponds to an item identified in the captured image with a threshold confidence. For example, a confidence metric may be generated based on the comparison of the image search result(s) and the captured image(s) to quantify the confidence in identifying the particular object. If the confidence does not meet a certain confidence threshold, the robotic computing device may decide to seek additional information about the particular object requested to be identified. However, if the confidence meets a certain threshold, the robotic computing device may perform the requested action of identifying the particular object or may delegate the task to another device.
[0046] In some embodiments, delegation of an action can be based on various factors, such as whether delegating the action reduces power consumption, alleviates network and / or processing bandwidth limitations on the device, completes the action faster, and / or other factors suitable for consideration when delegating actions between devices. For example, a robotic computing device may determine that an additional device (e.g., another robotic device or another smart device) is closer to a predicted location of an object. Based on this determination, the robotic computing device may determine that delegating the action of determining the object's location will consume less energy than at least performing the action itself.
[0047] If the robotic computing device determines to delegate one or more actions based on one or more different factors, method 300 may proceed from operation 304 to operation 314. Operation 314 may include the robotic computing device providing input therefrom to an additional computing device. Per the previous example, this input may be a communication (e.g., audio, visual, wireless, Bluetooth, etc.) between the robotic computing device and another computing device. The content of the communication may embody a request for the other computing device to navigate to and / or confirm the location of an object (e.g., Maggie's drone). Method 300 may proceed from operation 314 to operation 316. Operation 316 may include causing the additional computing device to perform an action to facilitate fulfillment of the request.
[0048] For example, once the robot computing device determines with a threshold confidence that the object's location can be determined by the robot computing device, the robot computing device may communicate with an additional computing device to determine the object's location. The communication from the robot computing device may be audio communication, visual communication, wireless communication, and / or any other communication that can be provided between devices. In some embodiments, the robot computing device may provide an audio output at a frequency higher or lower than that naturally detectable by humans. Alternatively or additionally, the communication from the robot computing device may be provided to the additional computing device as a communication via a wireless communication protocol (e.g., Bluetooth, Wi-Fi, LTE, and / or any other communication protocol). The communication may be processed by the additional computing device, such that the additional computing device may locate the object according to one or more different processes. For example, the additional computing device may be another robot computing device capable of navigating a different environmental region than the robot computing device upon receiving a communication from the robot computing device. Once the additional computing device determines the object's location, the additional computing device may communicate the object's location to the robot computing device.
[0049] If the robotic computing device determines not to delegate the action to another computing device in operation 304, method 300 may proceed from operation 304 to operation 306. Operation 306 may include determining whether to request additional information from another computing device and / or another user (with prior permission from the other user). In some embodiments, this determination may be based on whether the confidence level satisfies a threshold confidence level and / or one or more other thresholds. For example, if the magnitude of the confidence metric does not satisfy another confidence metric, the robotic computing device may determine to seek information from additional computing devices (e.g., moving from operation 306 to operation 308). Alternatively or additionally, the robotic computing device may rely on one or more heuristic processes and / or one or more trained machine learning models to determine whether to request additional information. For example, input from a user may be processed using one or more trained machine learning models to generate an embedding that can be compared to existing embeddings mapped to a latent space. If the embedding distance between the generated embedding and the existing embedding meets a threshold, method 300 may proceed from operation 306 to operation 312.
[0050] The operation 312 may include causing the robotic computing device to perform one or more actions (i.e., operations) to facilitate fulfillment of the request. For example, following the previous example, the robotic computing device may locate the object with a threshold confidence and indicate the location to the user. In some embodiments, the robotic computing device may indicate the location by providing an audible and / or visual output at one or more interfaces of the robotic computing device and / or another computing device. Alternatively or additionally, the robotic computing device may indicate the location of the object by suggesting movement to the object's location (e.g., "Okay, let's take you to Maggie's toy") and then move to the object's location.
[0051] If the robot computing device determines to request additional information from another computing device, the method 300 may proceed from operation 306 to operation 308. Operation 308 may include moving the robot computing device toward and communicating with the additional computing device. For example, the additional computing device may be a standalone computing device such as a display device or a speaker device. The robot computing device may select an additional computing device to communicate with based on a determination that the user has previously interacted with the additional computing device. For example, with prior permission from the user, the robot computing device may identify one or more devices with which the user has previously interacted and determine whether the one or more devices provide access to an automated assistant. If the robot computing device determines that a particular additional device provides access to an automated assistant, it may initiate communication with the automated assistant. For example, the robot computing device may identify a type of automated assistant accessible through the additional computing device and, based on the type of automated assistant, determine an invocation phrase for communicating with the automated assistant. For example, the robot computing device may use an application programming interface (API) and / or perform an internet search to identify the invocation phrase of the additional computing device.
[0052] Once the invocation phrase is identified, the robot computing device may generate natural language content including the invocation phrase and a specific request for the automated assistant. The invocation phrase may be, for example, "Assistant..." and the specific request may be, for example, "What does Maggie's drone look like?" The content of the specific request may be based on characteristics of the user input that may be uncertain to the robot computing device. The function identified as the basis for obtaining additional information may be selected based on one or more parameters and / or slot values of an action to be performed by the robot computing device to facilitate fulfillment of the request from the user. For example, the specific slot value may be based on an identifier of the object to be located. Thus, to generate the identifier, the robot computing device may request additional information about the object to be located from an additional computing device. Next, method 300 may proceed from operation 308 to operation 310.
[0053] Operation 310 may include receiving information regarding the request from the additional computing device. The additional information may be communicated via the same modality through which the robotic computing device communicated with the additional computing device. Alternatively or additionally, the additional computing device may communicate with the robotic computing device via a modality different from the modality through which the robotic computing device communicated with the additional computing device. For example, the robotic computing device may convey the request for additional information via an audio output to the additional computing device. In some embodiments, audio output from the robotic computing device may be provided to the additional computing device via beamforming and / or by selectively directing audio and / or other output to a particular additional computing device. In some embodiments, the audio output of the robotic computing device may be rendered to embody a voice that is acceptable to a speaker identification (“ID”) process utilized by the additional computing device with prior permission from the user. The voice may be selected by the user and / or based on prior interactions between the robotic computing device and the additional computing device. In some embodiments, the technique of beamforming output may be implemented using one or more microphones and / or antennas capable of detecting output provided by the robotic computing device. Based on the relative amplitude and / or phase of the signal detected at each microphone and / or antenna, each individual output interface (e.g., speaker, transmitter, etc.) may be adjusted to facilitate transmission to additional computing devices through constructive interference. In some embodiments, this beamforming technique may be utilized in response to the robotic computing device deciding to delegate a task (i.e., an action) to and / or seek additional information from another computing device.
[0054] In response to the audio output from the robotic computing device, the additional computing device may render an image on its display panel. The robotic computing device may then capture the image of the display panel and / or download the image via a wireless communication protocol (e.g., Bluetooth, Wi-Fi, etc.). Based on the additional information provided by the additional computing device, method 300 may then proceed from operation 310 and return to operation 304 to determine whether to delegate an action to another device, perform an action, and / or request further information from another device.
[0055] 4 is a block diagram 400 of an example computer system 410. Computer system 410 typically includes at least one processor 414 that communicates with multiple peripheral devices via a bus subsystem 412. These peripheral devices may include, for example, a storage subsystem 424 including memory 425 and a file storage subsystem 426, a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices allow a user to interact with computer system 410. Network interface subsystem 416 provides an interface to external networks and is coupled to corresponding interface devices in other computer systems.
[0056] The user interface input devices 422 may include a keyboard, a pointing device (e.g., a mouse, trackball, touchpad, graphic tablet, etc.), a scanner, a touchscreen integrated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and methods for inputting information into the computer system 410 or a communications network.
[0057] The user interface output devices 420 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or other mechanism for producing a visible image. The display subsystem may also provide non-visual displays, such as via an audio output device. In general, the use of the term "output device" is intended to encompass all possible types of devices and methods for outputting information from the computer system 410 to a user or to another machine or computer system.
[0058] The storage subsystem 424 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 424 may include logic to perform selected aspects of the method 300 and / or logic to implement one or more of the robotic computing device 104, the system 200, and / or any other applications, devices, apparatuses, and / or modules described herein.
[0059] These software modules typically execute on the processor 414 alone or in combination with other processors. The memory 425 used by the storage subsystem 424 may include multiple memories, such as a main random access memory (RAM) 430 for storing instructions and data during program execution, and a read-only memory (ROM) 432 in which fixed instructions are stored. The file storage subsystem 426 provides persistent storage of program files and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that perform the functions of a particular implementation may be stored by the file storage subsystem 426 within the storage subsystem 424, or may be stored on another machine accessible by the processor(s) 414.
[0060] Bus subsystem 412 provides a mechanism that allows the various components and subsystems of computer system 410 to communicate with each other as intended. Although bus subsystem 412 is shown schematically as a single bus, alternative embodiments of the bus subsystem may use multiple buses.
[0061] Computer system 410 can be of various types, such as a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Because computer and network capabilities vary, the description of computer system 410 shown in Figure 4 is intended as only one specific example to illustrate some embodiments. Many other configurations of computer system 410 are possible, having more or fewer components than the computer system shown in Figure 4.
[0062] Where the systems described herein collect or may utilize personal information about users (also sometimes referred to herein as “participants”), users may be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how they receive content from content servers that is more relevant to them. Additionally, certain data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be processed to prevent personal information from being identified to the user, or, when geographic location information (e.g., to the city, zip code, or state level) is obtained, the user's geographic location may be generalized to prevent the user's specific geographic location from being identified. Thus, users have control over how information about them is collected and used.
[0063] While several embodiments have been described and illustrated herein, various other means and / or structures can be utilized to perform the functions described herein and / or obtain the results and / or one or more advantages described herein, and each such variation and / or modification is deemed to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are for illustrative purposes only, and the actual parameters, dimensions, materials, and / or configurations will vary depending on the specific application(s) for which the present disclosure is utilized. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Accordingly, it should be understood that the above-described embodiments are presented by way of example only, and that, within the scope of the appended claims and their equivalents, embodiments may be practiced other than as specifically described and claimed. Embodiments of the present disclosure relate to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is within the scope of the present disclosure, unless such features, systems, articles, materials, kits, and / or methods are mutually inconsistent.
[0064] In some embodiments, a method executed by one or more processors is provided, the method including receiving, at the robot computing device, user input requesting the robot computing device to locate a particular object within an environment of the user and the robot computing device. The method may further include, in response to the user input, determining whether the robot computing device can locate the particular object within the environment with a threshold level of confidence. If the robot computing device is unable to locate the particular object within the environment with the threshold level of confidence, the method may further include moving the robot computing device within the environment toward another location closer to an additional computing device based on the robot computing device not locating the particular object with the threshold level of confidence. The method may further include generating a robotic input by the robot computing device and providing the robotic input as input to the additional computing device, the robotic input including content requesting information associated with the particular object from the additional computing device. The method may further include receiving, by the robot computing device, a response output from the other computing device in response to the robotic input, the response output characterizing particular information associated with the particular object. The method may further include causing the robot computing device to locate the particular object within the environment for the user based on the particular information.
[0065] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0066] In some embodiments, causing the robotic computing device to locate the particular object may include moving the robotic computing device to another location within the environment to facilitate locating the particular object.
[0067] In some embodiments, receiving a response output from the additional computing device may include capturing an image of a display panel of the additional computing device, the display panel rendering specific information to the robotic computing device when the image is captured.
[0068] In some embodiments, determining whether the robotic computing device can locate a particular object in the environment with a threshold confidence level includes processing home graph data that characterizes the locations of various features of the environment in which the user and the robotic computing device reside, and generating a confidence metric that characterizes the probability that a particular one of the various features of the environment corresponds to a particular object, and comparing the confidence metric to the threshold confidence level.
[0069] In some embodiments, the method may further include causing the robotic computing device to locate the particular object in the environment if the robotic computing device can locate the particular object in the environment with a threshold confidence.
[0070] In some embodiments, a method executed by one or more processors is provided, the method including receiving, at the robotic computing device, user input embodying a request for the robotic computing device to assist in performing an operation. The robotic computing device is present in an environment with a user and one or more other computing devices. The method may further include, in response to the user input, determining whether an additional computing device of the one or more other computing devices exhibits a state more suitable for initializing performance of the operation compared to a current state of the robotic computing device. The method may further include, if it is determined that the additional computing device of the one or more other computing devices exhibits a state more suitable for initializing performance of the operation, causing the robotic computing device to generate and provide as input to the additional computing device a robotic input. The robotic input includes content requesting the additional computing device to initialize performance of the operation. Providing the robotic input causes the additional computing device to initialize performance of the operation based on the robotic input.
[0071] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0072] In some embodiments, an additional computing device among the other computing device(s) is determined to exhibit a state that is more suitable for initiating execution of an operation based on (a) the additional computing device being connected to a utility power source, and (b) the current state of the robotic computing device being battery-powered.
[0073] In some embodiments, an additional computing device among the other computing device(s) is determined to be in a state more suitable for initiating execution of an operation based on the additional computing device indicating a signal strength that is greater than the signal strength indicated by the robotic computing device in its current state.
[0074] In some embodiments, an additional computing device among the other computing device(s) is determined to be indicative of a state that is more suitable for initiating execution of the operation based on the additional computing device indicating a processing bandwidth that is greater than the processing bandwidth indicated by the robot computing device in its current state.
[0075] In some embodiments, if it is determined that the additional computing device does not exhibit a state more suitable for initiating execution of the operation, the method further includes causing the robotic computing device to initiate execution of the operation.
[0076] In some embodiments, having the robotic input generated by the robotic computing device and provided as input to the additional computing device includes having the robotic computing device render audio output via an audio interface of the robotic computing device.
[0077] In some embodiments, causing the robotic computing device to generate the robotic input and provide it as input to the additional computing device includes determining a location of the additional computing device and causing the robotic computing device to render audio output toward the location of the additional computing device. In some aspects of these embodiments, causing the robotic computing device to render the audio output via an audio interface of the robotic computing device includes repositioning the robot and / or one or more components of the robot so that one or more speakers of the robot that emit the audio output are oriented toward the location of the additional computing device. Optionally, causing the robotic computing device to render the audio output via an audio interface of the robotic computing device includes rendering the audio output at a frequency above the human audible frequency range.
[0078] In some embodiments, a method executed by one or more processors is provided, the method including receiving, at a robotic computing device, user input embodying a request for assistance in performing an operation from the robotic computing device. The robotic computing device is present in an environment with a user and one or more other computing devices. The method may further include, in response to the user input, determining whether an additional computing device of the one or more other computing devices can perform the operation more effectively than the robotic computing device based on a location of the additional computing device. If the robotic computing device determines that the location of the additional computing device enables the additional computing device to perform the operation more effectively, the method may further include providing robotic input from the robotic computing device to the additional computing device to facilitate the additional computing device performing the operation. The robotic input causes the additional computing device to initiate performance of the operation and perform the request embodied in the user input.
[0079] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0080] In some embodiments, determining whether the location of the additional computing device allows the additional computing device to perform the operation more effectively than the robotic computing device includes determining that the location of the additional computing device is closer to a particular area of the environment than the robotic computing device, i.e., determining that the location of the additional computing device is closer to the particular area than the robot location of the robotic computing device.
[0081] In some embodiments, determining whether the location of the additional computing device allows the additional computing device to perform the operation more effectively than the robotic computing device includes determining that the location of the additional computing device is closer to another user in the environment than the robotic computing device, i.e., determining that the location of the additional computing device is closer to the other user than the robot location of the robotic computing device (i.e., that the robot location is farther from the other user than the location of the additional computing device). In these embodiments, fulfilling the request includes communicating with the other user.
[0082] In some embodiments, providing robotic input from the robotic computing device to the additional computing device to facilitate causing the additional computing device to perform the operation includes rendering an audible output at the robotic computing device embodying natural language content that instructs the additional computing device to perform the operation. In some aspects of these embodiments, the additional computing device provides access to an automated assistant, and the natural language content includes an invoke phrase that invokes the automated assistant. In some of the invoke phrase embodiments, the method further includes identifying, by the robotic computing device, a particular type of the additional computing device, and including, by the robotic computing device, the invoke phrase in the natural language content in response to identifying the invoke phrase, if any, for the particular type. Identifying the particular type of the additional computing device may optionally include processing image data captured by a camera of the robotic computing device and identifying the particular type based on processing the image data. In some additional or alternative aspects of these embodiments, the operation includes sending a message to another user, and the natural language content characterizes the message to be sent by the additional computing device to the other user.
Claims
1. A method executed by one or more processors, comprising: receiving user input at the robotic computing device requesting the robotic computing device to locate a particular object within an environment of the user and the robotic computing device; determining whether the robotic computing device can identify the location of the particular object within the environment with a threshold confidence in response to the user input; the robotic computing device is unable to determine the location of the particular object within the environment with the threshold confidence; based on the robotic computing device not identifying the location with the threshold confidence, moving the robotic computing device within the environment toward another location closer to an additional computing device; causing the robotic computing device to generate and provide a robotic input as input to the additional computing device, the robotic input including content requesting information associated with the particular object from the additional computing device; receiving, by the robotic computing device, a response output from the additional computing device in response to the robotic input, the response output characterizing particular information associated with the particular object; causing the robotic computing device to determine the location of the particular object within the environment for the user based on the particular information; A method comprising:
2. Having the robotic computing device identify the location of the particular object comprises: The method of claim 1 , comprising moving the robotic computing device to another location within the environment to facilitate locating the particular object.
3. Receiving the response output from the additional computing device includes: capturing an image of a display panel of the additional computing device; The method of claim 1 , wherein the display panel is rendering the specific information to the robotic computing device when the image is captured.
4. Determining whether the robotic computing device can identify the location of the particular object in the environment with the threshold confidence includes: processing home graph data that characterizes the positions of various features of the environment in which the user and the robotic computing device reside; generating a confidence metric based on processing the home graph data, the confidence metric characterizing the probability that a particular one of the various features of the environment corresponds to the particular object; comparing the confidence metric to the threshold confidence; The method of claim 1 , comprising:
5. If the robotic computing device can identify the location of the particular object in the environment with the threshold confidence, causing the robotic computing device to determine the location of the particular object within the environment; The method of claim 1 further comprising:
6. one or more processors; a memory storing instructions that, when executed, cause the one or more processors to perform a method according to any one of claims 1 to 5; A mobile robotic computing device comprising:
7. A computer program comprising instructions which, when executed, cause one or more processors to carry out the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Robot and robot system
JP2003345435A