Robotic computing device with adaptive user-interaction

The robotic computing device addresses the limitations of existing systems by autonomously navigating and communicating through natural language processing, enhancing its ability to assist users in location-based tasks.

JP2025090562AActive Publication Date: 2025-06-17GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025001773
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-23
Filing Date
2025-01-06
Publication Date
2025-06-17
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

Existing robotic computing devices lack the ability to autonomously navigate to various destinations and communicate information effectively among users without manual control, limiting their assistance in tasks that require location-based interactions.

Method used

The implementation of a robotic computing device that can navigate to users based on spoken utterances, determine likely locations of other users by accessing interaction data and home graph data, and communicate information through natural language processing and machine learning models.

Benefits of technology

Enables the robotic computing device to autonomously perform tasks, communicate effectively between users, and conserve resources by reducing the need for explicit user requests and manual control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090562000001_ABST
    Figure 2025090562000001_ABST
Patent Text Reader

Abstract

To provide a robotic computing device capable of performing various actions such as performing communication between users in a shared space in accordance with certain user preferences.SOLUTION: When interacting with a particular user, a robotic computing device can perform an operation at a preferred location relative to the particular user on the basis of an express or implied preference of that particular user. For instance, certain types of operations can be performed at a first location within a room, and other types of operations can be performed at a second location within the room. When an operation involves following or guiding a user, parameters for driving the robotic computing device can be selected on the basis of preferences of the user and / or a context in which the robotic computing device is interacting with the user (e.g., whether or not the context indicates some amount of urgency).SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Relates to a robotic computing device using adaptive user interaction.

Background Art

[0002] When a computing device enables interaction between an automatic assistant and a user, most computing devices cannot autonomously navigate to various destinations without being manually controlled by the user. This may limit the capabilities of some automatic assistants to provide assistance for some tasks that may involve navigating towards and / or away from the user. For example, a user who assigns the task of rendering a certain audio content to the user's automatic assistant may be restricted to locations where the user can perceive the audio content. This may result from the fact that audio content is often rendered via a stand-alone speaker device and / or other computing devices that must be manually placed by the user, which is often close to consent.

[0003] In some cases, when a user requests that certain actions be performed that require an automated assistant to move between geographical locations (e.g., between different rooms in a house), some tasks can be delegated to a device capable of performing that action. However, often such a device may not be capable of handling various actions. For example, a robotic vacuum cleaner may be able to initiate a default cleaning operation on request by the automated assistant, but may not be able to perform other cleaning-related actions with specificity. This can be the result of the automated assistant and / or the robotic vacuum cleaner not having a mechanism for converting the requests presented to the automated assistant by the user. This can be particularly inefficient when a robotic home device has some interfaces (e.g., speakers, sensors, etc.) that may be required to fulfill certain assistant requests, but the robotic home device does not have a mechanism for converting a particular assistant request into an executable operation that can be performed by the robotic home device. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0004] The implementations described in this specification relate to robot devices that can perform various tasks in which the robot device can be involved in navigating to a user for a particular purpose and / or communicating information among users. For example, the robot device can operate in a user's home and receive spoken utterances from the user such as "Robot, ask Eli if she's ready for school." In this situation, the user can be the person to whom the inquiry from the user is directed, who can be Eli's father. In response to the spoken utterance and with prior permission from a person in the home, the robot device can navigate from the living room where the robot device and the user were initially located and move towards Eli's room. In some implementations, the robot device can determine Eli's likely location by accessing data that can correlate several locations in the home to several names (e.g., "Eli's room"). Alternatively or additionally, the robot device can determine Eli's likely location based on a previous conversation between the robot device and Eli. For example, the other user, Eli, may have been located in the home office area during the most recent conversation between the robot device and Eli. Based on this determination and in response to the spoken utterance from the father, the robot device can navigate to the home office area to find Eli, i.e., the other user.

[0005] In some implementations, the robotic computing device can determine a likely location of another user, "Eli", by accessing interaction data that can indicate where the other user has recently interacted with another device. For example, just before the user provides the above-described verbal utterance, the other user may have interacted with a stand-alone display device in the user's home kitchen. Home graph data can provide a correlation between the stand-alone display device and the home kitchen, based at least on the stand-alone display device having a label in the home graph data, such as "kitchen display". Alternatively or additionally, semantic labels for rooms can be inferred from room characteristics, such as a room being labeled "kitchen" based on, for example, the room being the only room in a residence that has a dishwasher and a range (as determined while using one or more sensors and one or more object recognition techniques). In some implementations, the robotic computing device or the stand-alone display device can use one or more user verification techniques (e.g., voice verification, face recognition, etc.) to determine that the previous interaction involved the other user, Eli. Thus, using this data, the robotic computing device can determine that the other user has recently interacted with a particular device (e.g., the kitchen display) and that the particular device is located in the home kitchen. Based on this determination, the robotic computing device can respond to the above-described verbal utterance from the user by navigating towards the home kitchen to facilitate starting a question-and-answer session with the other user, Eli.

[0006] When the robot device identifies the location of the other user who is the subject of the verbal utterance, the robot device can issue an output via the output interface of the robot device (e.g., display interface, audio interface, etc.) with the prior permission of the other user. For example, the output issued to the other user can be a cordial audio output such as "Hi, Eli, your father is kind and wishes to know if you are ready for school." When the robot device issues an audio output and / or identifies Eli's location, the robot device can activate one or more input interfaces to receive input from Eli. For example, Eli may have heard the user asking the robot device to inquire whether Eli is ready for school, and as a result, can provide input to the robot device to respond before the robot device has an opportunity to render the audio output. In such a situation, since one or more microphones of the robot device will be preemptively activated before the other user arrives at the location, the robot device will be able to capture input from Eli. In other cases, when the other user, Eli, provides input after the robot device has rendered an audible output, the other user can provide a responsive input such as "I told dad I'm almost ready."

[0007] When the robot device receives an input in response from another user, the robot device can cause the input in response to be processed and can also begin to navigate back to the user who provided the first verbal utterance. In some implementations, the robot device can access a speech processing module, which can process audio data and / or text data using one or more trained machine learning models. For example, one or more trained machine learning models can include a transformer neural network and / or other language models that can be employed to convert the input in response into a meaningful output. For example, the robot device can cause audio input data corresponding to a verbal utterance from the user and / or another verbal utterance from another user to generate audio output data. The audio output data can be characterized by a truthful natural language output such as "Eli gently indicated that Eli is almost ready to go." In this way, a more natural type of Q&A can be created between the robot device and each user, as opposed to solely providing a verbatim recitation of what the responding user has said. This enables a clearer dialogue between the user and the user's robot device and can reduce instances where the robot device is asked to repeat what another user has said to the robot device. Some resources, such as battery life and processing bandwidth, can be conserved as a result.

[0008] In some implementations, the robotic computing device can assist the user in identifying a particular device and / or the location of a particular device in situations where the user cannot explicitly request the robotic computing device to find a particular device. In some cases, the robotic computing device can perform such an operation when rendering a notification that a particular device is not close enough to the user and / or the device is operating in silent mode and the user cannot confirm the response. For example, a cellular phone in a user's home may operate in silent mode when receiving an incoming phone call. The cellular phone can vibrate in silent mode, but the user may not be able to determine that the cellular phone is vibrating when the user and the cellular phone are in different rooms. However, the robotic computing device can receive a notification that the cellular phone is receiving an incoming call via a local area network (e.g., Wi-Fi) with prior permission from the user.

[0009] In response to receiving a notification, the robotic computing device can render an output such as "Sir, your phone is ringing silently." This output can be rendered when the robotic computing device is within a threshold distance of the user and / or after the robotic computing device has navigated towards the user in response to the notification. When the user can hear the output from the robotic computing device, the user can respond with a verbal utterance such as "Oh, thank you, I thought my phone was right here." This verbal utterance can be captured by the robotic computing device via the audio interface of the robotic computing device, converted into audio data, and the audio data can be processed at the robotic computing device and / or another specific computing device (such as a network device like a server) with the user's prior permission. The audio data can be processed to determine the user's intention to be assisted by the robotic computing device even though the user did not explicitly request the assistance of the robotic computing device.

[0010] For example, audio data and / or other data can be processed using one or more heuristic processes and / or one or more trained machine learning models (e.g., transformer neural network models, convolutional neural networks, recurrent neural networks, and / or other models). In some implementations, a robotic computing device, with prior permission from the user, can employ a neural network-based sequence classification model to determine whether the user presented a questioning tone and / or a certain amount of uncertainty. For example, audio data can be processed to generate a metric regarding the amount of uncertainty the user may be presenting with respect to a topic embodied in the user's spoken utterance and / or the output from the robotic computing device. When the metric meets a specific metric threshold, the audio data can be further processed to determine information and / or identify actions that can assist the user in analyzing the user's uncertainty and / or questions. For example, the action of identifying the location of a cellular phone can be determined by the robotic computing device to be useful for analyzing the detected uncertainty of the user.

[0011] Based on this determination, the robotic computing device can offer to perform an action by rendering another output such as "If you'd like, shall I take you to the phone?" In some implementations, the user can consent to the robotic computing device taking the user to the cellular phone by providing an explicit response input such as "Sure." Alternatively or additionally, the user can provide approval for the robotic computing device to take the user to the cellular phone by presenting body language and / or other characteristics indicating that the user wishes to be led to the cellular phone by the robotic computing device. For example, in response to the robotic computing device providing another output, the user can rise from where they are sitting and walk towards the robotic computing device. In some implementations, audio data and / or image data captured by the robotic computing device with the user's prior permission can be processed to determine whether the user is presenting an affirmative and / or approving response (i.e., another output) to the offer from the robotic computing device. In some implementations, the audio data and / or image data can be processed using one or more of the same or different trained machine learning models that were used to process the user's spoken utterances. For example, with the user's prior permission, multiple captured images can be processed to determine that the user's trajectory is towards the robotic computing device when the user is rising from the user's seat. Based on this determination, the robotic computing device can conclude that the user has presented an intention to be led to the cellular phone by the robotic computing device.

[0012] In some implementations, the speed or acceleration of a robotic computing device and / or the urgency with which the robotic computing device responds can be based on one or more characteristics of the context in which the robotic computing device initiated movement. For example, the type of notification received by the robotic computing device from another computing device indicates the urgency of the notification and can thus provide a basis for the speed at which the robotic computing device proceeds. For example, the robotic computing device can operate at a first speed when the robotic computing device is attempting to locate a ringing cellular phone. However, the robotic computing device can operate at a second speed lower than the first speed when the robotic computing device is attempting to locate the cellular phone in response to an incoming text message. Alternatively or additionally, the robotic computing device can establish a speed for moving to a particular location based on one or more different factors such as the urgency detected in the user's voice, the content of a particular notification, the source of a particular notification (e.g., the robotic computing device can move faster when a spouse sends a text message compared to when an acquaintance sends a text message), the time at which the notification is received, the application that is the source of the notification (e.g., the robotic computing device can move faster when receiving a delivery notification from a shopping application compared to a social media application that provides the notification), and / or any other source of data that can be a basis for the robotic computing device to move to different locations.

[0013] In some implementations, the robotic computing device can move to a particular location within a particular room depending on one or more characteristics of the operation being performed by the robotic computing device and / or the context in which the robotic computing device is performing the operation. For example, a user who is preparing dinner for guests can request that the robotic computing device play music in the kitchen of the user's home. In response to an explicit request to the robotic computing device (e.g., by providing an oral utterance such as "Play some music."), the robotic computing device can move to a particular location in the living room of the home to play music when the guests are in the home. In some implementations, this particular location can be learned by the robotic computing device and / or explicitly identified by the user (e.g., "Play music at this particular location in the living room when the guests are here.").

[0014] Moreover, in various implementations, this particular location can be specific to the operation being performed. For example, when performing the operation of "playing music", the robotic computing device can move to a first location in the room, and when performing the operation of "read me the news", the robotic computing device can move to a separate second location in the room, and when performing the operation of "streaming video from a smart camera", the robotic computing device can move to yet another separate third location in the room. In some implementations, the user can specify the location within the room for performing an operation, or type of operation, by providing a descriptive natural language input (e.g., "whenever you play videos, play them 1 meter in front of the north side of the couch."). Alternatively or additionally, the user can specify the location within the room for performing an operation, or type of operation, by moving themselves to the desired location (e.g., "whenever you tell me the news, tell me the news right here [the user walks to the desired location and stands there]."). The robotic computing device can then capture image data (and / or other sensor data), with prior permission from the user, and process the image data to determine the coordinates of where the user is standing and / or facing and / or the inferred semantic labels for them. Alternatively or additionally, the user can specify the location within the room for performing an operation, or type of operation, by interacting with the graphical user interface (GUI) of the robotic computing device or a separate device.For example, a first user can interact with a GUI to annotate a map of the first user's home to specify where some types of actions (e.g., audibly rendering news) should be performed. In some implementations, the map can be a graphical representation of a home or other structure with semantic labels generated by processing data from one or more sensors of a robotic computing device. Alternatively or additionally, a second user can also interact with an instance of the GUI to specify where some other types of actions (e.g., enabling a video call) should be performed.

[0015] In various implementations, preferences regarding locations for a robotic computing device to perform some actions and / or types of actions can be stored as user-specific preferences. For example, a first user can direct the robotic computing device to perform music-type actions at a first location in the living room at home, and a second user can direct the robotic computing device to perform music-type actions at a second location different from the first location in the living room. Thus, in response to the first user requesting that the robotic computing device "play music in the living room," the robotic computing device can verify that the first user is making the request and move to the first location within the living room according to the specified preference of the first user.

[0016] For example, before a guest arrives, the user can provide instructions for the robotic computing device to stay in the living room when performing any audio and / or video rendering. The user can provide these instructions via an oral utterance such as "When the guest arrives, please play only music next to the couch there." Optionally, the user can point to a location (e.g., when saying "...there..."), and the robotic computing device can capture the image data the user is pointing at (with prior permission from the user) to determine the exact location the user is pointing at. For example, the geographical layout data characterizing the user's home can be compared to the estimated trajectory of the user's pointing finger to determine the preferred "music" location the user is pointing at. Alternatively or additionally, the user can interact with a GUI to annotate a map of the user's home to specify where the user would prefer certain types of actions to be performed. Then, before the guest arrives, the robotic computing device can follow the user around the home (with or without playing music) as the user prepares for the guest's arrival. When the guest arrives, the robotic computing device can move to a location adjacent to the couch in the living room and either play music or wait for the user to explicitly request that the robotic computing device play music (e.g., "Play some music.").

[0017] In some implementations, this request from the user can be processed to generate preference data, and the preference data can be utilized when responding to subsequent requests from the user. For example, when the user has different guests over a subsequent weekend, the user can respond to a request to play music by moving to a location adjacent to the couch in the user's living room. In some implementations, user verification can be performed by the robotic computing device before implementing some preferences for some actions. For example, in response to the user requesting to play music, voice verification and / or face recognition can be performed. When the requesting user does not correspond to a user with specific location preferences for playing music, the robotic computing device can select a location based on one or more heuristic processes and / or one or more trained machine learning models. For example, a user who has not provided explicit preferences for a location for some actions may have some preferences inferred based on previous interactions between the user and the robotic computing device.

[0018] In some implementations, the location of the robotic computing device can be dynamic for some operations and / or for some user preferences. For example, a user can provide a request for the robotic computing device to issue an alarm after a certain duration (e.g., 10 minutes). In response, the robotic computing device can initially fail to follow the user but can initiate a "chase" operation after the duration before the alarm reaches a certain value. In some implementations, this certain value can be based on one or more characteristics of the alarm request, the reason for the requested alarm, the estimated relative location of the user with respect to the robotic computing device, and / or the context in which the user requested the alarm. In some implementations, a user can provide a request to the robotic computing device that causes the robotic computing device to chase (e.g., follow) the user without the user explicitly requesting that the robotic computing device follow the user. This behavior can be learned over time through interaction with the robotic computing device and / or based on explicit requests from the user. For example, a user can explicitly request that the robotic computing device stop chasing the user in some situations or stop chasing the user in other situations. Such cases can be used to train the robotic computing device to chase the user in some situations and / or not to chase the user in other situations without the user explicitly providing a request for the robotic computing device to chase the user, and can provide feedback data.

[0019] For example, a user can request that a robotic computing device play music, and in response, the robotic computing device can initialize a following operation and a music rendering operation. In some implementations, the following operation can be initialized in some contexts where the user is determined not to have another audio device nearby and / or is not walking towards the audio device. For example, when the user is walking along a corridor and provides a verbal utterance such as "Play some music", the robotic computing device can determine that the corridor has no existing audio devices. Alternatively or additionally, the robotic computing device can determine that the user is walking along the corridor in the direction of a room without an audio device. In response to the verbal utterance, the robotic computing device can initialize music rendering so that the user can listen to music while the user is walking into the room along the corridor, and can also initialize the following operation. Instead, when the room is determined to have an audio device, the robotic computing device can initialize playing music while the user is in the corridor, but when the user enters the room, can delegate playing music to the audio device in the room. Thereafter, when the user is in the room, the robotic computing device can optionally remain outside the room and abort the following operation, at least according to the user's determined preferences.

[0020] In some implementations, the characteristics of the following action may depend on the actions being performed by the robotic computing device and / or the type of actions being performed, and / or the specific user for whom the robotic computing device is providing requirements to perform the action. For example, when performing the action of playing music, the robotic computing device follows the user by a distance "x", but when enabling a video call or an audio call, it can follow the user at a different distance "y". Alternatively or additionally, the robotic computing device can follow the user at a specific speed according to the preferences of a specific user and / or the actions or types of actions being performed. In some implementations, the characteristics of the following action may be based on where the robotic computing device is located and / or whether the robotic computing device is located in a room with a specific inferred semantic label. For example, when in the "kitchen", the robotic computing device performs the following action according to a certain distance and / or speed, but when in the "garage", it can perform the following action according to a different distance and a different speed.

[0021] The above description is provided as an overview of some implementations of the present disclosure. Further descriptions of those implementations and other implementations are described in more detail below.

[0022] Other implementations may include a non - transitory computer - readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform one or more of the methods described above in this specification and / or described elsewhere, such as one or more of the methods. Other implementations may include one or more computer systems including one or more processors operable to execute stored instructions to perform one or more of the methods described above in this specification and / or described elsewhere, such as one or more of the methods.

[0023] It should be understood that all combinations of the above - described concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein.

Brief Description of the Drawings

[0024]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Figure 5

Figure 6

[0025] Figures 1A and 1B show FIGS. 100 and 120 in which a user 102 interacts with a robotic computing device 104 that can infer whether the user desires to be led to a device providing a notification. Alternatively or additionally, the robotic computing device can move to a device and / or a specific location according to drive operation parameters selected based on one or more characteristics of the context in which the notification is provided. For example, the robotic computing device 104 can determine that the user 102 has received a text message when the cellular phone, in a separate device such as a cellular phone, is operating in silent mode (e.g., vibration-only mode). Each of the cellular phone and the robotic computing device 104 can access a local area network, and the local area network can be wirelessly accessible to the robotic computing device 104 in the user 102's room 106. In some implementations, and with prior permission from the user 102, the robotic computing device 104 can determine whether the user 102 has confirmed a response to the text message and / or whether the user 102 is located in an area where the user 102 can detect the receipt of the text message. When the robotic computing device 104 predicts that the user 102 has not confirmed a response to the text message and / or is not located in an area where detection of the text message is possible, the robotic computing device 104 can render an audible output 108 such as "You have received a text message from Julian." Alternatively or additionally, the robotic computing device 104 can proactively render the audible output 108 when the cellular phone receives a text message.

[0026] In some implementations, the robotic computing device 104 can generate a prediction, with prior permission from the user 102, regarding whether the user 102 is aware of the location of the cellular phone. The prediction can be based on, for example, an oral utterance 110 indicating that the user 102 is not aware of the location of the cellular phone. The oral utterance 110 can include content such as "I don't even know where my phone is". Alternatively or additionally, image data and / or audio data can be processed, with prior permission from the user 102, to determine whether the user 102 knows the location of the cellular phone. For example, in response to the audible output 108, the user 102 can look around for the user 102's cellular phone (e.g., look to the left and right of the user 102), which can be an indication that the user 102 is not aware of the location of the user 102's cellular phone. Based on one or more of these contextual features, the robotic computing device 104 can determine that the user 102 is not aware of the location of the user 102's cellular phone and provide another audible output 112 such as "If you'd like, I can show you".

[0027] In some implementations, the robot computing device 104 can infer that the user 102 is interested in being led to the user 102's cellular phone without the user 102 providing an input that explicitly directs the robot computing device 104 to lead the user 102 all the way to the cellular phone. For example, the context feature 114 can include the user 102 rising from the user 102's seat after other audible outputs from the robot computing device 104. This feature 114 can be a positive indication that the user 102 intends to be led to the user 102's cellular phone by the robot computing device 104. For example, as shown in FIG. 120 of FIG. 1B, the robot computing device 104 can exit the room 106 and enter another room 126 to facilitate leading the user 102 to a device 124 such as the user 102's cellular phone. In this way, the user 102 does not necessarily have to engage in an explicit conversation with the user 102's assistant device to achieve some benefits. This can conserve the computing resources of the robot computing device 104 and / or other devices because less processing and memory can be consumed before initiating the performance of an operation.

[0028] Figures 2A, 2B, and 2C show FIGS. 200, 220, and 240 in which user 202 can communicate and interact with robot computing device 204 that can communicate among users at a learned preferred location. For example, user 202 can provide an oral utterance 208 such as "Go see if Jimmy has cleaned his room." The oral utterance can be directed to robot computing device 204, which may be located in room 206 with user 202. In response to detection of oral utterance 208, robot computing device 204 can render an audible output 210 such as "Okay, I'll go to his room and check." In some implementations, robot computing device 204 can determine that oral utterance 208 embodies one or more requests for robot computing device 204 to perform one or more operations. The one or more operations can include determining a location for the room of "Jimmy" and determining whether that room has been "cleaned."

[0029] In some implementations, the robot computing device 204 can determine the location of Jimmy's room by communicating with other computing devices that are located within a threshold distance of the robot computing device 204 and / or are connected to a common network with the robot computing device 204. For example, in response to determining a requested action, the robot computing device 204 can cause one or more devices inside and outside of the room 206 to render one or more different types of outputs (e.g., visual, audio, antenna, etc.) that can be detected by the robot computing device 204. The robot computing device 204 can determine, based on these outputs, that one or more of the devices correspond to Jimmy's room. Alternatively or additionally, the robot computing device 204 can determine the relative distance of the robot computing device 204 from one or more smart devices based on one or more signal metrics (e.g., signal quality) optionally detected within a time window or duration. In some implementations, user-specified names for certain devices can be "Jimmy's speaker" and / or "Jimmy's smart light", and thus can provide evidence that the location of a particular device corresponds to Jimmy's room. Alternatively or additionally, one or more signal metrics associated with communication between a particular device and the robot computing device 204 can indicate the relative distance and / or relative location of the particular device from the robot computing device 204.

[0030] In some implementations, such labels for a device can be identified via a smart home graph, an assistant application, and / or any other application or module that may be associated with the user 202 that is accessible to the robotic computing device 204. Alternatively or in addition, the robotic computing device 204 can determine a particular room where Jimmy was previously located and / or is currently located, with prior permission from one or more users. For example, the historical conversation data accessible to the robotic computing device 204 can indicate that most of the conversations between the robotic computing device 204 and another user (i.e., Jimmy) occurred in a different room 226, as shown in FIG. 2B. Based on this determination and the absence of conflicting data (e.g., data indicating that room 226 is not Jimmy's room), the robotic computing device 204 can navigate to room 226 to fulfill a request from the user 202.

[0031] When the robotic computing device 204 arrives in another room 226, the robotic computing device 204 can collect data about the other room 226 by using one or more interfaces of the robotic computing device 204 and / or interacting with one or more other devices within the other room 226. For example, the robotic computing device 204 can utilize one or more cameras to capture one or more images of the other room 226. The one or more images can be processed using one or more trained machine learning models to determine whether the other room 226 should be classified as "clean". In some implementations, the robotic computing device 204 can position itself in the other room 226 based on the preferences of the user 202 and / or another user 222. For example, a first location preference for the robotic computing device 204 within the other room 226 can correspond to a location where the robotic computing device 204 should be positioned to collect data. Alternatively or additionally, a second location preference for the robotic computing device 204 within the other room 226 can correspond to another location where the robotic computing device 204 should be positioned to interact with another user 222.

[0032] In some implementations, the robot computing device 204 can infer location preferences based on frequency at which a user engages with the robot computing device 204 at several locations within a room. Alternatively or additionally, the robot computing device 204 can determine the user's location preferences based on explicit instructions from the user. For example, the user can provide explicit requests for the robot computing device 204 to enable visual output and / or video calls only at some preferred distances and / or only at specific portions of the room. Alternatively or additionally, the user can provide explicit requests for the robot computing device 204 to enable visual output and / or video calls only at other distances and / or only at other portions of the room. These instructions can be received and generated to be preference data, which can then be utilized by the robot computing device 204 when interacting with one or more users.

[0033] According to the example above, the robot computing device 204 can enter another room 226 and optionally render an output 228 to another user 222 to facilitate fulfilling a request from the user 202. For example, the other output 228 can be a request about information that can assist the robot computing device 204 in fulfilling the request from the user 202, such as "Did you clean the room?" In response, the other user 222 (i.e., Jimmy) can provide a verbal utterance 230 such as "Yes, I'm working on it." This information can optionally be used in combination with other data generated by the robot computing device 204 to fulfill the request from the user 202. For example, upon receiving this information and / or data, the robot computing device 204 can navigate to the user 202 as shown in FIG. 2C.

[0034] In some implementations, the robotic computing device 204 can navigate to a specific location 246 within the room 206 based on the preferences of the user 202. For example, the user 202 can prefer that the robotic computing device 204 provide its response at a specific location 246 while the robotic computing device 204 is entering the room 206 to provide an audible response. Alternatively or additionally, the user 202 can have a history of guiding the robotic computing device 204 to a specific location 246 by providing a fetch command (e.g., "Come here.") in combination with a pointing gesture (e.g., a movement pointing towards the specific location 246). Based on these historical interactions, the robotic computing device 204 can determine that the user 202 prefers to hear a response from the robotic computing device 204 at a specific location 246 within the room 206. In some implementations, the preferred distance and / or the preferred location can vary from room to room and / or for different operations and / or different types of outputs. For example, the user 202 can prefer that the robotic computing device 204 provide "news" from the end of the bed when the robotic computing device 204 is in the user 202's bedroom, while the robotic computing device 204 can enable an audio phone call from the side of the bed when the robotic computing device 204 is in the user 202's bedroom.

[0035] When the robotic computing device 204 arrives at a particular location 246 to fulfill a request, the robotic computing device 204 can render an audible output 242 such as "The room looks clean. Jimmy says he's working on it." This audible output 242 can be generated using one or more language models (e.g., recurrent neural network, transformer network model, etc.). In some implementations, the audible output 242 may contain different content than that provided by another user 222, but still conveys a similar conclusion. Alternatively or additionally, the audible output 242 can include content based on a response from the user 222 and data generated using one or more interfaces (e.g., camera) of the robotic computing device 204. In this way, the rendered content from the robotic computing device 204 can embody natural language content that characterizes the acquired data and information to facilitate fulfilling the request from the user 202.

[0036] Figure 3 shows a system 300 that can operate a robotic computing device to enable communication between users and move to a specific location to complete an operation without an explicit request from the user. The system 300 can include a computing device 302 that can be a robotic computing device, which includes one or more applications to enable the robotic computing device to interface with the user. For example, the robotic computing device can include an automatic assistant 304. The automatic assistant 304 can operate as part of an assistant application provided on one or more other computing devices and / or server devices. The user can interact with the automatic assistant 304 via an assistant interface 320, which can be a microphone, a camera, a touch screen display, a user interface, and / or any other device capable of providing an interface between the user and the application. For example, the user can initialize the automatic assistant 304 by providing linguistic, text, and / or graphical input to the assistant interface 320 to cause the automatic assistant 304 to initialize one or more actions (e.g., provide data, control peripheral devices, access agents, generate input and / or output, etc.). Alternatively, the automatic assistant 304 can be initialized based on processing context data 336 using one or more trained machine learning models.

[0037] Context data 336 can characterize one or more features of the environment accessible to the automatic assistant 304 and / or one or more features of the user predicted to intend to interact with the automatic assistant 304. The computing device 302 can include a display device, which can be a display panel including a touch interface for receiving touch inputs and / or gestures to enable the user to control the application 334 of the computing device 302 via the touch interface. In some implementations, the computing device 302 can be without a display device and thereby can provide an audible user interface output without providing a graphical user interface output. Further, the computing device 302 can provide a user interface, such as a microphone, for receiving oral natural language input from the user. In some implementations, the computing device 302 can include a touch interface and may be without a camera, but optionally can include one or more other sensors.

[0038] Computing device 302 and / or other third-party client devices may communicate with a server device via a network, such as the Internet. Additionally, computing device 302 and any other computing devices may communicate with each other via a local area network (LAN), such as a Wi-Fi network. Computing device 302 may offload computing tasks to the server device to conserve computing resources in computing device 302. For example, the server device may host the automatic assistant 304, and / or computing device 302 may send inputs received at one or more assistant interfaces 320 to the server device. However, in some implementations, the automatic assistant 304 may be hosted on computing device 302, and various processes that may be associated with the automatic assistant operation may be performed on computing device 302.

[0039] In various implementations, all or fewer than all aspects of the automatic assistant 304 may be implemented on the computing device 302. In some of those implementations, aspects of the automatic assistant 304 may be implemented via the computing device 302 and interface with a server device that can implement other aspects of the automatic assistant 304. The server device may optionally service multiple users and the users' associated auxiliary applications via multiple threads. In implementations where all or fewer than all aspects of the automatic assistant 304 are implemented via the computing device 302, the automatic assistant 304 may be an application that is separate from (e.g., installed "on top of") the operating system of the computing device 302, or alternatively, may be implemented directly by the operating system of the computing device 302 (e.g., considered an application of the operating system but integral with the operating system).

[0040] In some implementations, the automatic assistant 304 can include an input processing engine 306, and the input processing engine 306 can employ multiple different modules to process input and / or output for the computing device 302 and / or the server device. For example, the input processing engine 306 can include an audio processing engine 308, and the audio processing engine 308 can process audio data received at the assistant interface 320 to identify text embodied in the audio data. The audio data can be transmitted from the computing device 302 to the server device, for example, to conserve computing resources at the computing device 302. Additionally or alternatively, the audio data can be processed solely at the computing device 302.

[0041] The process for converting audio data to text can include a speech recognition algorithm that can employ a neural network and / or a statistical model for identifying groups of audio data corresponding to words or phrases. The text converted from the audio data is parsed by a data parsing engine 310 and made available to the automatic assistant 304 as text data that can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, the output data provided by the data parsing engine 310 can be provided to a parameter engine 312 to determine whether the user provided input corresponding to specific intents, actions, and / or routines that can be performed by the automatic assistant 304 and / or an application or agent that can be accessed via the automatic assistant 304. For example, the assistant data 338 can be stored in the server device and / or the computing device 302 and can include data that defines one or more actions that can be performed by the automatic assistant 304 and the parameters required to perform those actions. The parameter engine 312 can generate one or more parameters for intents, actions, and / or slot values and provide the one or more parameters to an output generation engine 314. The output generation engine 314 can use the one or more parameters to communicate with an assistant interface 320 for providing output to the user and / or with one or more applications 334 for providing output to one or more applications 334.

[0042] In some implementations, the automatic assistant 304 can be an application that can be installed "on top of" the operating system of the computing device 302, and / or can itself form part (or all) of the operating system of the computing device 302. The automatic assistant application includes, and / or has access to, on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition can be implemented using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model stored locally in the computing device 302. On-device speech recognition generates recognized text for any spoken utterances (if any) present in the audio data. Also, for example, on-device natural language understanding (NLU) can be implemented using an on-device NLU module that processes the recognized text generated using on-device speech recognition and optionally context data to generate NLU data.

[0043] NLU data can include an intent corresponding to an oral utterance and optionally parameters (e.g., slot values) for that intent. On-device fulfillment can be implemented using an on-device fulfillment module that utilizes NLU data (from on-device NLU) and optionally other local data to determine the actions to take to analyze the intent of the oral utterance (and optionally the parameters for that intent). This can include determining local and / or remote responses (e.g., answers) to the oral utterance, interactions with locally installed applications for execution based on the oral utterance, commands for transmission to Internet of Things (IoT) devices based on the oral utterance (either directly or via a corresponding remote system), and / or other analytical actions for execution based on the oral utterance. On-device fulfillment can then initiate the local and / or remote implementation / execution of the actions determined to analyze the oral utterance.

[0044] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment can be utilized at least selectively. For example, the recognized text can be transmitted at least selectively to a remote assistant component for remote NLU and / or remote fulfillment. For example, the recognized text can be transmitted, optionally, in parallel with on-device implementation or in response to a failure of on-device NLU and / or on-device fulfillment, for remote implementation. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution can be prioritized at least by the latency reduction they provide when analyzing the oral utterance (by not requiring a client-server round trip to analyze the oral utterance). Additionally, on-device functionality can be the only functionality available in situations where there is no or limited network connectivity.

[0045] In some implementations, computing device 302 can include one or more applications 334 that can be provided by a third - party entity different from the entity that provided computing device 302 and / or the automatic assistant 304. The application state engine of the automatic assistant 304 and / or the computing device 302 can access application data 330 to determine one or more actions that can be performed by one or more applications 334, as well as the state of each application of the one or more applications 334 and / or the state of each device associated with the computing device 302. The device state engine of the automatic assistant 304 and / or the computing device 302 can access device data 332 to determine one or more actions that can be performed by the computing device 302 and / or one or more devices associated with the computing device 302. Further, the application data 330 and / or any other data (e.g., device data 332) can be accessed by the automatic assistant 304 to generate context data 336, which can characterize the context in which a particular application 334 and / or device is running, and / or the context in which a particular user is accessing the computing device 302, application 334, and / or any other device or module.

[0046] While one or more applications 334 are running on computing device 302, device data 332 can characterize the current operating state of each application 334 running on computing device 302. Further, application data 330 can characterize one or more features of the running application 334, such as the content of one or more graphical user interfaces being rendered in the direction of one or more applications 334. Alternatively or in addition, application data 330 can characterize an action schema, which can be updated by each application and / or by the automatic assistant 304 based on the current operating status of each application. Alternatively or in addition, the one or more action schemas for one or more applications 334 can remain static, but can be accessed by the application state engine to determine suitable actions to be initialized via the automatic assistant 304.

[0047] The computing device 302 can further include an assistant invocation engine 322, which can use one or more trained machine learning models to process application data 330, device data 332, context data 336, and / or any other data accessible to the computing device 302. The assistant invocation engine 322 can process this data to determine whether the user should be waiting to explicitly speak a wake phrase to invoke the assistant 304, or whether that data can be considered to indicate the user's intent to attempt to invoke the assistant without the user having to explicitly speak a wake phrase. For example, one or more trained machine learning models can be trained using examples of training data based on scenarios where a user is in an environment where multiple devices and / or applications are exhibiting various operating states. Examples of training data can be generated to capture training data that characterizes contexts where the user invokes the assistant and other contexts where the user does not invoke the assistant. When one or more trained machine learning models are trained according to these examples of training data, the assistant invocation engine 322 can, based on the characteristics of the context and / or environment, cause the assistant 304 to detect or limit the detection of a spoken wake phrase from the user.

[0048] In some implementations, the assistant call engine 322 can process data generated using one or more assistant interfaces 320 to determine whether the user is expressing an intention to benefit from the operation of the automatic assistant 304 and / or the robotic computing device. For example, data captured via one or more assistant interfaces 320 can be processed by the assistant call engine 322 to determine whether the user's verbal utterances, non-verbal gestures, disfluency, and / or other movements can be regarded as a call to the robotic computing device. Alternatively or additionally, the data can be processed to determine the context in which the user provided such input and / or movement. Based on the determined context and / or expression of the user, the robotic computing device can determine whether the user has an intention to enable the robotic computing device to perform a particular operation. The operation can be, for example, providing information to the user and / or guiding the user to a particular location even though the user does not provide an explicit request for the robotic computing device to do so.

[0049] In some implementations, system 300 can include a drive parameter engine 316 that can determine one or more parameters for moving a robotic computing device based on some contexts and / or some data. For example, data providing a basis for a particular action to be performed by a robotic computing device can be processed by the drive parameter engine 316 to determine how the robotic computing device should move when performing the particular action. For example, application data 330, device data 332, and / or context data 336 can be processed to determine whether there is an urgency and / or time limit associated with a particular request from a user. In some implementations, this can be determined using one or more heuristic processes and / or one or more trained machine learning models. Alternatively or additionally, data associated with a request can be processed by the drive parameter engine 316 to generate an embedding that can be mapped to a latent space, and the distance in the latent space to a certain point and / or area can indicate whether the request is urgent. This processing by the drive parameter engine 316 can be utilized to determine drive parameters such as speed, acceleration, travel time, power limits, and / or any other parameters that can be associated with driving a robotic device.

[0050] For example, application data 330 can indicate that the user has requested to be led to the location of the device that provided the emergency notification. The application data 330 can be processed by the drive parameter engine 316 to determine that the application exhibits a particular status and that the particular status is predicted to be particularly urgent with respect to other notifications and / or other application statuses. In some implementations, this determination can be based on whether the application status has a time quality (e.g., the phone is ringing, an important contact is calling, and thus there is only a certain amount of time to answer the phone call). Based on this determination, the drive parameter engine 316 can identify speed parameters for controlling one or more motors of the robotic computing device when fulfilling the request to be led to the location of the device. In some implementations, the system 300 can include a layout detection engine 318, and the layout detection engine 318 can enable the robotic computing device to determine the relative location and / or other features of the room within the space or structure in which the robotic computing device is located. For example, the layout detection engine 318 can be utilized when the robotic computing device attempts to identify the location of the device, the user, the room, and / or the features of the space and / or structure in response to input from the user.

[0051] In some implementations, the layout detection engine 318 can cause other devices to provide an output to assist in determining the current location of the robotic computing device relative to other parts of the space in which the robotic computing device is located. For example, the layout detection engine 318 can process the application data 330 to determine that a device in a particular room has a user-defined label (e.g., "laundry room speaker") that can indicate the name of the particular room (e.g., "laundry room"). When the robotic computing device is directed to enter a particular room, the layout detection engine 318 can cause the device to provide an output (e.g., illuminate a light or display, render audio, transmit an antenna signal, etc.). The layout detection engine 318 can identify one or more different characteristics (e.g., signal metrics) of the output (e.g., signal quality, amplitude, magnitude, audio frequency, optical frequency, etc.) received from one or more different devices within a time window or duration to determine the relative location of the particular room compared to the current location of the robotic computing device.

[0052] For example, a robotic computing device can determine whether the robotic computing device is collocated with one or more devices having semantic labels associated with a room specified or inferred from user requests. Based on this determination, the layout detection engine 318 can determine how the robotic computing device should move from its current location to the location of the desired room. In some implementations, locations within a structure (e.g., a home or a company) can be mapped using public knowledge graph analysis and / or personal knowledge graph analysis. The public knowledge graph can be generated based on prior interactions between one or more users, people, and one or more other applications. Alternatively or additionally, a private knowledge graph can be generated based on prior interactions between a user and one or more applications (e.g., an assistant application and / or an IoT application).

[0053] In some implementations, the system 300 can include a location preference engine 326 that can perform several operations, interact with a particular user, and / or in some cases, determine a preferred location of the robotic computing device for positioning the robotic computing device. The location preference engine 326 can process data using one or more heuristic processes and / or trained machine learning models to determine whether the current location of the robotic computing device is suitable for fulfilling a request from the user. In some implementations, the preferred location can be requested directly by the user and / or inferred from one or more prior interactions with one or more users. For example, the user can explicitly or implicitly request that the robotic computing device perform several operations (e.g., enable a phone call or video call) at a first location within a room and other operations (e.g., play music) at a second location within the room. These preferred locations can be different for different users and / or for different rooms.

[0054] For example, when a first user receives a voice call via a robotic computing device, the robotic computing device may navigate to a first location in the room. However, when a second user receives a voice call, the robotic computing device may navigate to a second location in the room so that the second user can receive the voice call via the robotic computing device. In some implementations, the location preference engine 326 may be able to determine a dynamic location for the robotic computing device. For example, a first user may prefer that the robotic computing device follow the first user by a distance x (where x is any distance value) when playing music. However, a second user may prefer that the robotic computing device follow the second user by a different distance y (where y is any distance value) when rendering news reports and conference calls.

[0055] FIG. 4 shows a method 400 for operating a robotic computing device to paraphrase user input and / or user responses that can be relayed by the robotic computing device from a first user to a second user. The robotic computing device can relay messages between locations at different speeds according to the type of input explicitly provided to the robotic computing device and / or inferred by the robotic computing device. Method 400 can be implemented by one or more computing devices, applications, and / or any other device or module that can be associated with an automated assistant. Method 400 can include an operation 402 of determining whether verbal utterances and / or other user input have been provided to the robotic computing device. The robotic computing device can be a computing device that can move to various locations within a user's home and connect to one or more different networks to send and receive data. When a verbal utterance or other user input is detected at the robotic computing device, method 400 can proceed to operation 404. In other cases, the robotic computing device can continue to determine whether the user has provided input.

[0056] Action 404 can include determining whether the user has directed the robotic computing device to perform an action that may require locomotion. The action that may require locomotion can include relaying a message to another user who may not be located within a distance at which the robotic computing device can emit audible and / or visual audio. For example, an oral utterance from a first user can include a request for the robotic computing device to provide an audible message to a second user located in a different room than the first user. The oral utterance from the first user can be, for example, "Tell Phoenix that I need to leave since Sherry just parked in front of the house." This oral utterance can embody a request for the robotic computing device to perform at least an action of moving to the location of the second user (e.g., "Phoenix") and another action of providing a message to the second user (e.g., "Sherry just parked in front of the house, so Mark is about to leave now."). When it is determined that a request for the robotic computing device to locomote to a different location has been received, method 400 can proceed from action 404 to action 406. In other cases, method 400 can proceed from action 404 to action 410.

[0057] Operation 406 can include determining whether a request from a user has increased relative importance. The increased relative importance can refer to the priority or severity of the request relative to other requests presented by the user and / or one or more other users. For example, in some implementations, voice characteristics of an oral utterance (such as cadence, words per second, etc.) can be identified and processed to determine whether the oral utterance is intended to have increased relative importance. In some implementations, an operation request corresponding to a time event can indicate the relative importance of the request. For example, a request associated with a situation that has a particular probability of not changing within at least a threshold duration can be considered to have no increased relative importance. However, a different request associated with a situation that has a particular probability of changing within at least a threshold duration can be considered to have increased relative importance.

[0058] When an oral utterance "I told Phoenix that since Sherry just parked in front of the house, I need to leave" is received by a robotic computing device, the oral utterance can be determined to embody a request of increased relative importance. This determination can be based at least in part on one or more heuristic processes and / or one or more trained machine learning models. For example, the robotic computing device and / or other computing devices can determine that the request embodied in the oral utterance characterizes a time event (such as the user needing to leave) and / or is provided in a tone indicating urgency (such as the cadence of the user's voice like "... I... need to... leave..."). Based on the historical conversation between the user and the robotic computing device and / or one or more other applications, the robotic computing device can determine that this tone is different from the user's general tone and indicates a sense of urgency.

[0059] When it is determined that the user has provided a request with increased relative importance, method 400 can proceed to operation 408. In other cases, method 400 can proceed to operation 412. Operation 408 can include causing the robotic computing device to proceed according to a first drive parameter. For example, the first drive parameter can include, without limitation, the amount of acceleration, the amount of speed, and / or the amount of energy utilized to fulfill the request from the user. When the first drive parameter is selected, the amount of energy consumed and / or the amount of speed of the robotic computing device can be selected to be increased relative to other values selected for requests of lower relative importance. For example, operation 412 can include causing the robotic computing device to proceed according to a second drive parameter that can be different from the first drive parameter. The second drive parameter can characterize a speed setting and / or an acceleration setting that can be lower than the setting corresponding to the first drive parameter. When the first drive parameter and / or the second drive parameter is employed to control the proceeding operation of the robotic computing device, method 400 can proceed from operation 408 or operation 412 to operation 410.

[0060] Operation 410 can include causing a robotic computing device to perform the requested operation. Operation 410 can be performed when the user arrives at the destination and / or on the way to the corresponding destination. The requested operation can be, for example, rendering an output to another user, identifying the location of another computing device, retrieving information available at the destination, and / or any other operation that the computing device can be requested to perform. To facilitate the above examples, the robotic computing device can render an audible output to a second user (i.e., Phoenix) such as "Mark is about to leave now since Sherry just parked in front of the house." In some implementations, the output from the robotic computing device can be different from the verbal utterance from the first user but can convey the information provided by the first user. Method 400 can proceed from operation 410 to operation 414 that determines whether it is predicted that the first user, the second user, or another user will provide an additional request to the robotic computing device.

[0061] A robotic computing device can determine whether it is predicted that a user will provide an additional request to the robotic computing device based on one or more direct and / or indirect gestures performed by the user. For example, the robotic computing device can determine that, with prior permission from the user, the user stared at the device, directed the user's voice towards the device, moved towards the device, and / or performed a gesture indicating an interest in providing an additional request to the robotic computing device. To facilitate the above example, operation 414 can be performed with respect to the first user and / or the second user. For example, when the robotic computing device reaches the second user and provides an audible output, the robotic computing device can determine whether it is predicted that the second user will provide an input to the robotic computing device, with prior permission from the second user. Alternatively or additionally, operation 414 can be performed when the robotic computing device returns to the first user after providing an audible output to the second user. When it is predicted that the user will provide additional input to the robotic computing device, method 400 can proceed to operation 416. In other cases, method 400 can return to operation 402 to determine whether an input has been provided to the robotic computing device.

[0062] Operation 416 can include causing the user to follow the robotic computing device or, in some cases, moving with the user, according to the predicted type of input. For example, when the robotic computing device predicts that a second user will provide an input in response to an audible output, the robotic computing device can move with the second user. In some implementations, the robotic computing device can move with the second user for a duration corresponding to the predicted type of input and / or a confidence score for the input prediction. For example, when the robotic computing device predicts a current input with a first confidence score, the robotic computing device can follow the user for a first duration. However, when the robotic computing device predicts another current input with a second confidence score that is greater than the first confidence score, the robotic computing device can follow the user for a second duration that is longer than the first duration. Thereafter, method 400 can return to operation 402 to determine whether the user has provided an input to the robotic computing device.

[0063] FIG. 5 shows a method 500 for operating a robotic computing device to autonomously assign semantic labels to one or more areas within a space or structure occupied by one or more users and / or robotic computing devices. The semantic labels can be assigned to several locations to facilitate ensuring that the robotic computing device performs some actions in some areas according to user preferences. Method 500 can be implemented by any computing device, application, and / or any device or module capable of interacting with the robotic computing device. Method 500 can include an operation 502 of determining whether an area of the space occupied by the robotic computing device is not associated with a semantic label. For example, the robotic computing device can operate in a home occupied by one or more users (e.g., a mother user and a daughter user), and the home can include various different areas (e.g., different parts such as a living room, a kitchen, an office, a bedroom, etc.). Operation 502 can be initialized when the robotic computing device is located in the kitchen of the home, before or in response to the user providing an input to the robotic computing device. In this way, each semantic label can assist the robotic computing device in performing specific actions that can be more effectively implemented using information related to locations in and / or near the robotic computing device. For example, when the robotic computing device is directed by a first user to communicate with a second user within the home, the robotic computing device can access the semantic labels for different parts of the home to predict the location of the second user and optionally, subsequently, the location of the first user, with the prior permission of the second user.

[0064] When the robotic computing device determines that the space occupied by or near the robotic computing device is not associated with a semantic label, method 500 can proceed from operation 502 to operation 504. In other cases, the robotic computing device can continue to determine whether the space in or near the robotic computing device is associated with a semantic label. Operation 504 can include causing a first set of smart devices to emit one or more first outputs during a first time window. The first set of smart devices can be selected to emit one or more first outputs based on the first set of smart devices being assigned labels with relevant content (e.g., "kitchen counter speaker", "refrigerator smart display", etc.). Alternatively or additionally, the first set of smart devices can include some devices located at one or more specific locations on a map generated over time using data from one or more sensors of the robotic computing device. For example, when the robotic computing device moves to different locations within a first user's home, the robotic computing device can capture (with prior permission from the user within the home) data characterizing the locations of some devices within the home. This data can be utilized to correlate these devices with locations on a mapping generated by the robotic computing device and / or one or more other devices.

[0065] Method 500 can proceed from operation 504 to operation 506, which can include causing a second set of smart devices to emit one or more second outputs during a second time window. In some implementations, the one or more first outputs and the one or more second outputs can include audio output and / or visual output. For example, the one or more first outputs can include one or more characteristics that are the same as or different from one or more other characteristics of the one or more second outputs. For example, the one or more first outputs can embody one or more frequencies (e.g., audio and / or visual) that are different from one or more other frequencies (e.g., audio and / or visual) embodied by the one or more second outputs. In some implementations, the second set of smart devices can be selected to emit one or more second outputs based on a determination that a robotic computing device is located in a room and / or area of space where the second set of smart devices is different from the first set of smart devices.

[0066] Method 500 can proceed to operation 508, which can include processing sensor data generated by one or more sensors of a robotic computing device. In some implementations, the first time window and the second time window can be durations that at least partially overlap or do not overlap. For example, sensor data captured during the first time window can have the same or different timestamps as other sensor data captured during the second time window. In some implementations, sensor data captured by the robotic computing device and / or one or more other computing devices can be processed to determine whether a first set of smart devices and / or a second set of smart devices are collocated with the robotic computing device. For example, the magnitude of a detected output during the first time window can be compared to the magnitude of another detected output during the second time window. This comparison can indicate whether the first set of smart devices or the second set of smart devices are collocated in or near the area occupied by the robotic computing device. Alternatively or additionally, one or more characteristics of the detected output can be determined and compared to a first set of characteristics and / or a second set of characteristics. When the first set of characteristics is detected as being embodied in the output detected by the robotic computing device, the robotic computing device can determine that the first set of smart devices is collocated in or near the robotic computing device. Alternatively or additionally, when the second set of characteristics is detected as being embodied in the output detected by the robotic computing device and optionally when there is no second set of characteristics, the robotic computing device can determine that the second set of smart devices is collocated in or near the robotic computing device.

[0067] In some implementations, the detected characteristics can include the magnitude of the output, the frequency of the output, a change in the magnitude of the output (e.g., compared to a set magnitude), a change or modulation in the frequency of the output (e.g., compared to a set frequency), and / or any other characteristic that can be associated with the rendered output. In some implementations, the audio and / or visual output can be rendered at one or more frequencies that are not visually detectable and / or not audibly detectable by a natural human and / or a human unaided by non-intrinsic features. In some implementations, the output rendered by a set of devices can include a combination of one or more different outputs (e.g., outputs rendered by different interface modalities) that can be distinguished from other outputs rendered by another set of devices.

[0068] When it is determined that a robotic computing device is collocated with a particular set of smart devices (e.g., a first set and / or a second set), method 500 can proceed from operation 510 to operation 512. Operation 512 can include generating a semantic label for an area within the space currently occupied by the robotic computing device. In some implementations, the set of smart devices can include, but is not limited to, smart lights, smart televisions, smart thermostats, smart speakers, and / or any other device that can be associated with a user and controlled via a separate device and / or application. In some implementations, the semantic label can be based on an existing descriptor of a room within a space or structure, where the descriptor was previously assigned to the detected set of smart devices. For example, the detected set of smart devices can be assigned a descriptor (e.g., "Sam's room [device type]") by the user in response to an explicit user input to an assistant and / or other device or application. Alternatively or additionally, the detected set of smart devices can be assigned a descriptor that is generated based on processing data using one or more heuristic processes and / or one or more machine learning models. When an area occupied by the robotic computing device is assigned a semantic label, the semantic label can be stored in relation to a location on a generated map that can be utilized by the robotic computing device to move between locations within the space or structure.

[0069] When a user provides an oral input that is synonymous with, for example, a semantic label, a robotic computing device can correlate the oral input with an area to which a semantic label is assigned. In this way, the robotic computing device can utilize existing data and / or other information for mapping the user's home without the user having to explicitly identify areas within the home to the robotic computing device. This enables the robotic computing device to fulfill several requirements with less information being explicitly provided by the user, thereby conserving the computing resources of the robotic computing device. For example, when a semantic label (e.g., "Sherry's room") is assigned to an area of a map, the robotic computing device can, in some instances, navigate to that area when a semantic label (or a term and / or part synonymous with the semantic label) is identified in a natural language input provided to the robotic computing device (e.g., "Go tell Sherry that dinner is ready"). When the robotic computing device determines that none of the sets of smart devices are collocated with the robotic computing device, method 500 can proceed from operation 510 to operation 514. Operation 514 can include moving the robotic computing device to different areas of the space (e.g., the home) occupied by the user and / or the robotic computing device and optionally performing operation 502.

[0070] FIG. 6 is a block diagram 600 of an exemplary computer system 610. The computer system 610 generally includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices can include, for example, a storage subsystem 624 including a memory 626 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with the computer system 610. The network interface subsystem 616 provides an interface to an external network and is coupled to a corresponding interface device in other computer systems.

[0071] The user interface input device 622 can include a keyboard, a mouse, a trackball, a pointing device such as a touchpad or a graphics tablet, a scanner, a touch screen incorporated in a display, a voice recognition system, an audio input device such as a microphone, and / or other types of input devices. Generally, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 610 or onto a communication network.

[0072] The user interface output device 620 can include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visual image. The display subsystem can also provide a non-visual display, such as via an audio output device. Generally, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 610 to a user or to another machine or computer system.

[0073] The memory subsystem 624 stores programming and data configurations that provide some or all of the functionality of the modules described herein. For example, the memory subsystem 624 can include logic for implementing selected aspects of method 400, method 500, and / or for implementing one or more of system 300, robotic computing device 104, robotic computing device 204, the automated assistant, and / or any other application, device, apparatus, and / or module described herein.

[0074] These software modules are generally executed by the processor 614 alone or in combination with other processors. The memory 626 used in the memory subsystem 624 can include several memories, such as a main random access memory (RAM) 630 for storing instructions and data during program execution, and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide a persistent storage device for program and data files and can include a hard disk drive, a floppy disk drive, a CD-ROM drive, an optical drive, or a removable media cartridge along with associated removable media. Modules implementing the functionality of some implementations can be stored by the file storage subsystem 626 in the memory subsystem 624 or in other machines accessible by the processor 614.

[0075] The bus subsystem 612 provides a mechanism for enabling the various components and subsystems of the computer system 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem can use multiple buses.

[0076] The computer system 610 can be of different types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computer system 610 shown in FIG. 6 is to be considered only a specific example for illustrating some implementations. Numerous other configurations of the computer system 610 are possible that have more or fewer components than the computer system shown in FIG. 6.

[0077] In situations where the systems described in this specification may collect or use personal information about a user (or often referred to herein as a "participant"), the user may be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographical location), or to control whether and / or how to receive content from a content server that may be relevant to the user. Also, certain data may be handled in one or more ways before being stored or used so that personally identifiable information is removed. For example, the user's identifying information may be handled so that personally identifiable information cannot be determined about the user, or the user's geographical location may be generalized, in which case the geographical location information is obtained such that the user's specific geographical location cannot be determined (up to city, zip code, or state level, etc.). Thus, the user may have control over how information is collected and / or used about the user.

[0078] Although several implementations have been described and illustrated herein, various other means and / or structures may be utilized to perform the functions and / or to obtain the results and / or one or more of the advantages described herein, and each such variation and / or modification is to be regarded as being within the scope of the implementations described herein. Those skilled in the art will come to recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the above-described implementations are presented by way of example only and it is to be understood that implementations may be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods do not mutually conflict.

[0079] In some implementations, a method implemented by one or more processors is provided, including the step of a robot computing device determining that a user has uttered a verbal utterance indicating that the user is unsure of the location of a particular computing device. The verbal utterance does not embody an explicit request for the robot computing device to identify the location of the particular computing device. The method can further include the step of the robot computing device causing an output interface of the robot computing device to provide an indication to the user that the robot computing device is capable of determining the location of the particular computing device. The method can further include the step of the robot computing device processing input data from one or more input interfaces of the robot computing device to facilitate determining whether the user is willing to allow the robot computing device to guide the user towards the location of the particular computing device. The method can further include the step of causing the robot computing device to communicate with the particular computing device and move the robot computing device towards the relative location of the particular computing device to facilitate estimating the relative location of the particular computing device with respect to the robot computing device when the robot computing device determines that the user is willing to allow the robot computing device to guide the user towards the location of the particular computing device.

[0080] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.

[0081] In some implementations, to facilitate a user in determining whether they have the intention to allow a robotic computing device to guide the user towards the location of a particular computing device, the step of processing input data includes the step of processing image data indicative of the user's movement towards the robotic computing device. In some implementations, the input data does not include audio data characterizing an explicit request from the user for the robotic computing device to determine the relative location of a particular computing device. In some implementations, the step of moving the robotic computing device towards the relative location of a particular computing device includes moving the robotic computing device towards the relative location of a particular computing device at a speed selected based on the status of an application accessible via the particular computing device. In some of those implementations, the application includes a voice call application, and the status of the application indicates that the user has missed a call from a particular contact. In some implementations, the step of causing the robotic computing device to communicate with a particular computing device to facilitate estimating the relative location of the particular computing device with respect to the robotic computing device is the step of determining a signal metric based on the communication between the robotic computing device and the particular computing device, the signal metric indicating the relative distance of the particular computing device from the robotic computing device. In some of those implementations, the signal metric includes the audio amplitude of an audio output rendered by the particular computing device.

[0082] In some implementations, a method implemented by one or more processors is provided that includes receiving, by a robotic computing device, an oral utterance from a first user located in a space with the robotic computing device and a second user. The method can further include determining, based on the oral utterance, that the first user has directed the robotic computing device to communicate with the second user. The second user is located at a second user location that is different from the first user location of the first user. The method can further include, in response to the oral utterance, causing the robotic computing device to move to the second user location and render an output for the second user. The output embodies a natural language question based on the oral utterance from the first user. The method can further include receiving, by the robotic computing device, a responsive input from the second user. The responsive input embodies natural language content responsive to the natural language question embodied in the output from the robotic computing device. The method can further include, after the robotic computing device provides the output for the second user, causing the robotic computing device to move to the first user location and render another output for the first user. The other output characterizes the responsive input from the second user and embodies other natural language content different from the natural language content embodied in the responsive input from the second user.

[0083] These and other implementations of the techniques disclosed herein can optionally include one or more of the following features.

[0084] In some implementations, for the robot computing device, the step of moving to the second user location includes the step of determining a location preference associated with the second user. In some versions of those implementations, for the robot computing device, the step of moving to the second user location includes the step of moving the robot computing device to a specific location corresponding to the preferred location indicated by the location preference. The location preference can indicate a preferred location for the robot computing device when the robot computing device communicates with the second user. In some of those versions, the preferred location indicates a preferred distance of the robot computing device from the second user, and the specific location is at least the preferred distance away from the second location of the second user. In some implementations, for the robot computing device, the step of moving to the first user location includes the step of determining a location preference associated with the first user and the step of moving the robot computing device to a specific location corresponding to the preferred location indicated by the location preference. The location preference can indicate a preferred location for the robot computing device when the robot computing device renders a specific type of output for the first user. For example, the specific type of output can include an audible output provided via an audio output interface of the robot computing device or a visual output provided via a display interface of the robot computing device. For example, the specific type of output can be an audible output with content characterizing a message from another user.

[0085] In some implementations, a method implemented by one or more processors is provided, including the step of determining that a user has requested that a robotic computing device perform an operation in a particular room located in a space that includes a plurality of different rooms. The method can further include the step of causing the robotic computing device to provide one or more respective outputs detectable by the robotic computing device to one or more devices in one or more of the plurality of different rooms. The method can further include the step of determining, based on the one or more respective outputs, whether the current location of the robotic computing device corresponds to the particular room. The method can further include the step of causing the robotic computing device to move to the particular room and causing the robotic computing device to perform the operation when the robotic computing device is located in the particular room, based on the fact that the current location of the robotic computing device does not correspond to the particular room when the current location of the robotic computing device does not correspond to the particular room.

[0086] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.

[0087] In some implementations, for the step of moving the robotic computing device to a specific room, in order to perform a specific type of operation corresponding to the above operation, the step includes determining that a specific part of the specific room is selected by the user, and the step of moving the robotic computing device to the specific part of the specific room. In some implementations, for the step of moving the robotic computing device to a specific room, in order to perform a specific type of operation not corresponding to the above operation, the step includes determining that a specific part of the specific room is selected by the user, and the step of moving the robotic computing device to a different part of the specific room. In some implementations, the method can further include, when the current location of the robotic computing device corresponds to a specific room, causing the robotic computing device to identify the current part of the current room that is the preferred part for performing the above operation within the current room where the robotic computing device is located. In some of those implementations, the step of causing the robotic computing device to identify the current part of the current room that is the preferred part of the room for performing the above operation includes determining, based on the user requesting the above operation, that the user previously requested the robotic computing device to perform a specific type of operation corresponding to the above operation in the preferred part of the room. In some implementations, the method can further include, when the current location of the robotic computing device does not correspond to a specific room, rendering an output that requests the user to confirm that the current location of the robotic computing device is approved for performing the above operation before the robotic computing device performs the above operation.In some implementations, the method can further include identifying, for a robot computing device, a relative distance for following a user while the user moves to another portion of a particular room within the current room in which the robot computing device is located, when the current location of the robot computing device corresponds to the particular room.

[0088] In some implementations, a method implemented by one or more processors is provided, based on a map generated by a mobile robot computing device and at least partially based on sensor observations of the mobile robot computing device, the mobile robot computing device determines that it is currently located within a particular area of a structure (e.g., within a room of a home). The method can further include causing a first subset of smart devices to each emit one or more first outputs while the mobile robot computing device is located within the particular area, and causing a second subset of smart devices to each emit one or more second outputs. The one or more first outputs are audible and / or visual, and the one or more first outputs are caused to be emitted during a first time window and / or with one or more first characteristics in response to each of the first subset of smart devices being assigned a first semantic label in a home graph. The one or more second outputs are audible and / or visual, and the one or more second outputs are caused to be emitted during a second time window and / or with one or more second characteristics in response to each of the second subset of smart devices being assigned a second semantic label in a home graph. The method can further include obtaining sensor data during the emission of the one or more first outputs and the one or more second outputs. The sensor data is generated by one or more sensors of the mobile robot computing device. The method can further include determining, based on the analysis of the sensor data, that a first subset of smart devices is collocated with the robot in the particular area.The step of determining that a first subset of smart devices is collocated with a robot in a particular area is based on (1) the detected output being within a first time window and / or matching one or more first characteristics, as indicated by the analysis, and / or (2) the magnitude of the detected output being within a first time window and / or matching one or more first characteristics. The method can further include the step of assigning an inferred semantic label to the particular area in response to a determination that a first subset of smart devices is collocated with a robot in a given room. The inferred semantic label can be the same as or derived from the first semantic label assigned to the first subset of smart devices in the home graph.

[0089] These and other implementations of the techniques disclosed herein can optionally include one or more of the following features.

[0090] In some implementations, one or more first outputs are emitted during a first time window, and one or more second outputs are emitted during a second time window. In some versions of those implementations, the step of determining that a first subset of smart devices is collocated with a robot in a particular area includes the step of determining that a detected output occurs during the first time window and the step of determining that there is no detected output that occurs during the second time window. In some additional or alternative versions of those implementations, the step of determining that a first subset of smart devices is collocated with a robot in a particular area includes the step of determining that the magnitude of the detected output that occurs during the first time window is greater than the additional magnitude of additional detected outputs that occur during the second time window. In some implementations, one or more first outputs have a first characteristic and one or more second outputs have a second characteristic. In some versions of those implementations, the step of determining that a first subset of smart devices is collocated with a robot in a particular area includes the step of determining that a detected output matches the first characteristic and the step of determining that there is no detected output that matches the second characteristic. In some of those versions, one or more first characteristics include a first frequency and one or more second characteristics comprise a second frequency. For example, the first output can include a visual output, the first frequency can be a first visual frequency, the second output can include a second visual output, and the second frequency can be a second visual frequency. In some implementations, one or more first outputs have a first characteristic, one or more second outputs have a second characteristic, and the step of determining that a first subset of smart devices is collocated with a robot in a particular area includes the step of determining that the magnitude of the first characteristic in the detected output is greater than the additional magnitude of the second characteristic in the detected output. In some versions of those implementations, one or more first characteristics include a first frequency and one or more second characteristics include a second frequency.For example, the first output can include an audible output, the first frequency is a first audible frequency outside the range of human hearing, the second output can include an audible output, and the second frequency is a second audible frequency outside the range of human hearing.

[0091] In some implementations, the first subset of smart devices includes stand-alone automatic assistant devices, and one or more of the first outputs include a first audible output via the hardware speakers of the automatic assistant device. In some implementations, the first subset of smart devices includes stand-alone automatic assistant devices, and one or more of the first outputs include a first visual output via the hardware display of the automatic assistant device or via the light-emitting diodes of the automatic assistant device. In some implementations, the first subset of smart devices includes smart lights, smart televisions, and / or smart thermostats. In some implementations, the first semantic label in the home graph is the first descriptor of the first room in the structure, previously assigned to the first subset of smart devices based on a first explicit user input, and / or the second semantic label in the home graph is the second descriptor of the second room in the structure, previously assigned to the second subset of smart devices based on a second explicit user input. In some implementations, the step of assigning an inferred semantic label to a particular area includes automatically assigning the inferred semantic label to a particular area in a map for use by a mobile robot device. In some versions of those implementations, the method can further include using the inferred semantic label when controlling the navigation of the mobile robot device, after automatically assigning the inferred semantic label to a particular area in a map for use by the mobile robot device.In some of those versions, when controlling the navigation of a mobile robot device, the step of using the inferred semantic label includes determining that one or more terms of the verbal input match the inferred semantic label based on processing the verbal input detected at one or more microphones of the mobile robot device, and navigating the robot to a specific area based on determining that one or more terms match the inferred semantic label and based on the inferred semantic label being assigned to a specific area in the map. In some implementations, the step of assigning the inferred semantic label to a specific area includes proposing to the user in a graphical user interface that the inferred semantic label be assigned to a specific area in the map for use by the mobile robot device, and assigning the inferred semantic label to a specific area in the map for use by the mobile robot device in response to receiving an affirmative user interface input from the user in response to the proposal.

Description of the Signs

[0092] 102 User 104 Robot Computing Device 106 Room 108 Audible Output 110 Verbal Utterance 112 Another Audible Output 114 Feature 124 Device 126 Another Room 202 User 204 Robot Computing Device 206 Room 208 Verbal Utterance 209 Audible Output 222 Other User, User 226 Another Room, Room 228 Output, Other Output 230 Oral speech 242 Audible output 246 Specific location 300 System 302 Computing device 304 Automatic assistant 306 Input processing engine 308 Audio processing engine 310 Data syntax analysis engine 312 Parameter engine 314 Output generation engine 316 Drive parameter engine 318 Layout detection engine 320 Assistant interface 322 Assistant call engine 326 Location preference engine 330 Application data 332 Device data 334 Application 336 Context data 338 Assistant data 610 Computer system 612 Bus subsystem 614 Processor 616 Network interface subsystem 620 User interface output device 622 User interface input device 624 Memory subsystem 626 Memory, file storage subsystem 630 Random access memory (RAM) 632 Read-only memory (ROM)

Claims

1. 1. A method implemented by one or more processors, the method comprising: determining, based on a map generated by the mobile robotic computing device and based at least in part on sensor observations of the mobile robotic computing device, that the mobile robotic computing device is currently located within a particular area of ​​a structure; While the mobile robotic computing device is located within the particular area, causing a first subset of smart devices to each emit one or more first outputs, the one or more first outputs being audible and / or visual, the one or more first outputs being emitted during a first time window and / or with one or more first characteristics in response to the first subset of smart devices each being assigned a first semantic label in a home graph; causing a second subset of smart devices to each emit one or more second outputs, the one or more second outputs being audible and / or visual, the one or more second outputs being emitted during a second time window and / or with one or more second characteristics in response to the second subset of smart devices each being assigned a second semantic label in the home graph; acquiring sensor data during the emission of the one or more first outputs and the one or more second outputs, the sensor data being generated by one or more sensors of the mobile robot computing device; determining, based on the analysis of the sensor data, that the first subset of smart devices are co-located with a robot in the particular area, the determining that the first subset of smart devices are co-located with the robot in the particular area comprising: the analysis indicating detected outputs during the first time window and / or matching the one or more first characteristics; and / or a magnitude of the detected output during the first time window and / or matching the one or more first characteristics; and In response to determining that the first subset of smart devices are co-located with the robot in a given room, assigning an inferred semantic label to the particular area, the inferred semantic label being the same as or derived from the first semantic label assigned to the first subset of smart devices in the home graph; A method comprising:

2. determining that the one or more first outputs are emitted during the first time window and the one or more second outputs are emitted during the second time window and that the first subset of smart devices are co-located with the robot in the particular area, comprising: determining that the detected output occurs during the first time window; and determining that no detected output occurs during the second time window.

2. The method of claim 1, comprising:

3. determining that the one or more first outputs are emitted during the first time window and the one or more second outputs are emitted during the second time window and that the first subset of smart devices are co-located with the robot in the particular area, comprising: determining that the magnitude of the detected output occurring during the first time window is greater than an additional magnitude of an additional detected output occurring during the second time window; 2. The method of claim 1, comprising:

4. determining that the one or more first outputs have the first characteristic and the one or more second outputs have the second characteristic and that the first subset of smart devices are co-located with the robot in the particular area, comprising: determining that the detected output matches the first characteristic; and determining that no detected output matches the second characteristic.

3. The method of claim 1 or claim 2, comprising:

5. The method of claim 4 , wherein the one or more first characteristics comprise a first frequency and the one or more second characteristics comprise a second frequency.

6. the first output comprises a visual output and the first frequency is a first visual frequency; the second output comprises a visual output and the second frequency is a second visual frequency. The method of claim 5.

7. determining that the one or more first outputs have the first characteristic and the one or more second outputs have the second characteristic and that the first subset of smart devices are co-located with the robot in the particular area, comprising: determining that a magnitude of the first characteristic in the detected output is greater than an additional magnitude of the second characteristic in the detected output; The method of claim 1 or claim 3, comprising:

8. The method of claim 7 , wherein the one or more first characteristics comprise a first frequency and the one or more second characteristics comprise a second frequency.

9. the first output comprises an audible output, and the first frequency is a first audible frequency that is outside the range of human hearing; the second output comprises an audible output and the second frequency is a second audible frequency outside the range of human hearing. The method of claim 8.

10. 10. The method of claim 1, wherein the first subset of smart devices comprises a standalone automated assistant device, and the one or more first outputs comprise a first audible output via a hardware speaker of the standalone automated assistant device.

11. 11. The method of claim 1, wherein the first subset of smart devices comprises a standalone automated assistant device, and the one or more first outputs comprise a first visual output via a hardware display of the automated assistant device or via a light emitting diode of the automated assistant device.

12. 12. The method of claim 1, wherein the first subset of smart devices comprises a smart light, a smart television, or a smart thermostat.

13. the first semantic label in the home graph is a first descriptor of a first room in a structure that was previously assigned to the first subset of smart devices based on a first explicit user input; the second semantic label in the home graph is a second descriptor of a second room within a structure previously assigned to the second subset of smart devices based on a second explicit user input; 13. The method according to any one of claims 1 to 12.

14. assigning the inferred semantic label to the particular area, automatically assigning the inferred semantic labels to the particular areas in the map for use by the mobile robotic device.

14. The method of any one of claims 1 to 13, comprising:

15. after automatically assigning the inferred semantic labels to the particular areas in the map for use by the mobile robotic device; using the inferred semantic labels in controlling navigation of the mobile robotic device.

15. The method of claim 14, further comprising:

16. using the inferred semantic labels in controlling navigation of the mobile robotic device, determining, based on processing a verbal input detected at one or more microphones of the mobile robotic device, that one or more terms of the verbal input match the inferred semantic labels; causing the robot to navigate to the particular area based on determining that the one or more terms match the inferred semantic label and based on assigning the inferred semantic label to the particular area in the map; 16. The method of claim 15, comprising:

17. The step of assigning the inferred semantic label to the particular area comprises: suggesting to a user in a graphical user interface that the inferred semantic label be assigned to the particular area in the map for use by the mobile robotic device; in response to receiving a positive user interface input from the user responsive to the suggestion, assigning the inferred semantic label to the particular area in the map for use by the mobile robotic device; 14. The method of any one of claims 1 to 13, comprising:

18. 1. A method implemented by one or more processors, the method comprising: determining, by the robotic computing device, that a user has uttered a verbal utterance indicating that said user is unsure of the location of a particular computing device; the verbal utterance does not embody an explicit request for the robotic computing device to identify the location of the particular computing device; causing an output interface of the robotic computing device to provide an indication to the user that the robotic computing device is capable of determining the location of the particular computing device; processing input data from one or more input interfaces of the robotic computing device to facilitate determining, by the robotic computing device, whether the user is willing to allow the robotic computing device to guide the user towards the location of the particular computing device; when the robotic computing device determines that the user is willing to allow the robotic computing device to guide the user toward the location of the particular computing device, causing the robotic computing device to communicate with the particular computing device to facilitate estimating a relative location of the particular computing device with respect to the robotic computing device; causing the robotic computing device to move towards the relative location of the particular computing device; A method comprising:

19. The step of causing the robotic computing device to move towards the relative location of the particular computing device comprises: moving the robotic computing device toward the relative location of the particular computing device at a speed selected based on a status of applications accessible via the particular computing device.

20. The method of claim 18, comprising:

20. 20. The method of claim 19, wherein the application includes a voice call application, and the status of the application indicates that the user has missed a call from a particular contact.

21. causing the robotic computing device to communicate with the particular computing device to facilitate estimating the relative location of the particular computing device with respect to the robotic computing device, determining a signal metric based on communication between the robot computing device and the particular computing device, the signal metric indicating a relative distance of the particular computing device from the robotic computing device.

21. The method of claim 20, comprising:

22. The method of claim 21 , wherein the signal metric comprises an audio amplitude of an audio output rendered by the particular computing device.

23. processing the input data to facilitate determining whether the user is willing to allow the robotic computing device to guide the user towards the location of the particular computing device, processing image data indicative of the user's movements towards the robotic computing device; 23. The method of any one of claims 18 to 22, comprising:

24. 24. The method of claim 18, wherein the input data is absent audio data characterizing an explicit request from the user for the robotic computing device to determine the relative location of the particular computing device.

25. 1. A method implemented by one or more processors, the method comprising: determining, at a robotic computing device, that a user has requested that the robotic computing device perform an operation in a particular room located in a space that includes a plurality of different rooms; causing, by the robotic computing device, one or more devices in one or more of the plurality of different rooms to provide one or more respective outputs detectable by the robotic computing device; determining whether a current location of the robotic computing device corresponds to the particular room based on the one or more respective outputs; when the current location of the robotic computing device does not correspond to the particular room; causing the robotic computing device to move to the particular room based on the current location of the robotic computing device not corresponding to the particular room; causing the robotic computing device to perform the action when the robotic computing device is located in the particular room; A method comprising:

26. The step of causing the robotic computing device to move to the particular room comprises: determining that a particular portion of the particular room is preferred by the user for performing a particular type of action corresponding to the action; causing the robotic computing device to move to the particular portion of the particular room; 26. The method of claim 25, comprising:

27. The step of causing the robotic computing device to move to the particular room comprises: determining that a particular portion of the particular room is preferred by the user for performing a particular type of action that does not correspond to the action; causing the robotic computing device to move to a different portion of the particular room; 26. The method of claim 25, comprising:

28. when the current location of the robotic computing device corresponds to the particular room; causing the robotic computing device to identify, within a current room in which the robotic computing device is located, a portion of the current room that is a preferred portion for performing the action; 26. The method of claim 25, further comprising:

29. Having the robotic computing device identify the portion of the current room that is the preferred portion of the room for performing the action comprises: determining, based on the user requesting the action, that the user has previously requested the robot computing device to perform a particular type of action corresponding to the action in the preferred portion of the room; 29. The method of any one of claims 25 to 28, comprising:

30. when the current location of the robotic computing device does not correspond to the particular room; causing the robotic computing device to render an output requesting the user to confirm that the current location of the robotic computing device is approved for performing the action before the robotic computing device performs the action.

30. The method of any one of claims 25 to 29, further comprising:

31. when the current location of the robotic computing device corresponds to the particular room; having the robotic computing device identify a relative distance within a current room in which the robotic computing device is located to follow the user when performing the action while the user moves to another part of the particular room.

31. The method of any one of claims 25 to 30, further comprising:

32. 1. A method implemented by one or more processors, the method comprising: receiving, by a robotic computing device, a verbal utterance from a first user located in a space with the robotic computing device and a second user; determining, based on the verbal utterance, that the first user has directed the robotic computing device to communicate with the second user, the second user is located at a second user location that is different from a first user location of the first user; in response to the verbal utterance, causing the robotic computing device to move to the second user location and render an output for the second user, the output embodying a natural language question based on the verbal utterance from the first user; receiving, by the robotic computing device, a responsive input from the second user, the responsive input embodied in natural language content responsive to the natural language question embodied in the output from the robotic computing device; after the robotic computing device has provided the output for the second user, causing the robotic computing device to move to the first user location and render another output for the first user, the other output characterizing the responsive input from the second user and embodying other natural language content that differs from the natural language content embodied in the responsive input from the second user; A method comprising:

33. The step of causing the robotic computing device to move to the second user location comprises: determining location preferences associated with the second user, the location preference indicating a preferred location for the robotic computing device when the robotic computing device communicates with the second user; causing the robotic computing device to move to a particular location corresponding to the preferred location indicated by the location preference; 33. The method of claim 32, comprising:

34. 34. The method of claim 33, wherein the preferred location indicates a preferred distance of the robotic computing device from the second user, and the particular location is at least the preferred distance away from the second location of the second user.

35. The step of causing the robotic computing device to move to the first user location comprises: determining location preferences associated with the first user, the location preferences indicating preferred locations for the robotic computing device when the robotic computing device renders a particular type of output for the first user; causing the robotic computing device to move to a particular location corresponding to the preferred location indicated by the location preference; 35. The method of any one of claims 32 to 34, comprising:

36. 36. The method of any one of claims 32 to 35, wherein the particular type of output comprises an audible output provided via an audio output interface of the robotic computing device, or a visual output provided via a display interface of the robotic computing device.

37. 37. The method of claim 36, wherein the particular type of output is the audible output with content that characterizes a message from another user.

38. 38. A computer program comprising instructions which, when executed by one or more processors of a computing system, cause the computing system to perform a method according to any one of claims 1 to 37.

39. 38. A system comprising one or more computing devices configured to carry out the method of any one of claims 1 to 37.

40. 40. The system of claim 39, wherein the one or more computing devices comprise a mobile robotic computing device.

41. 38. A computer-readable storage medium storing instructions executable by one or more processors of a computing system to implement a method according to any one of claims 1 to 37.

Citation Information

Patent Citations

  • Monitoring device, monitoring device control method, server, server control method, monitoring system, and control program

    JP2015011463A

  • Cloud robotics system, information processor, program, and method for controlling or supporting robot in cloud robotics system

    JP2017047519A

  • System and method for controlling an autonomous mobile robot

    JP2019525273A

  • Method, system, and device for creating a map of device status, controlling, and displaying device status - Patents.com

    JP2020522156A

  • Cleaning robot, home monitoring apparatus, and method for controlling the cleaning robot

    US20140324271A1