Motorized computing device that autonomously adjusts device location and / or interface orientation in accordance with automated assistance requests
The mobile computing device addresses the challenge of users accessing information while engaged in tasks by navigating to the user and adjusting its display panel's viewing angle based on the user's location, enhancing user convenience and conserving resources.
Patent Information
- Application Number
- JP2023158377
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2039-04-29
AI Technical Summary
Users face challenges in accessing information from computing devices when they are engaged in tasks that prevent them from comfortably or safely navigating to the device, such as working in a different room or being in a location where they cannot see or hear the content rendered by the device.
A mobile computing device that can selectively navigate to a user and adjust the viewing angle of its display panel based on the user's relative location, using motors and sensors to determine the user's position and render content accordingly, while conserving power by only navigating when necessary.
The mobile computing device effectively renders content to users in various locations without requiring them to pause their tasks or navigate to the device, thereby improving user convenience and conserving computational resources.
Smart Images

Figure 0007689553000001 
Figure 0007689553000002 
Figure 0007689553000003
Abstract
Description
[Background technology]
[0001] Humans may engage in human-computer dialogue with interactive software applications, referred to herein as "automated assistants" (also referred to as "digital agents," "chatbots," "interactive personal assistants," "intelligent personal assistants," "conversational agents," etc.). For example, a human (who may be referred to as a "user" when interacting with an automated assistant) may provide commands and / or requests using oral natural language input (i.e., speech) and / or by providing textual (e.g., typed) natural language input, which may in some cases be converted to text and then processed. While the use of automated assistants can allow easier access to information and more convenient means for controlling peripheral devices, perceiving display and / or audio content may be difficult in some situations.
[0002] For example, when a user is engrossed in some tasks in one room of the house and desires to obtain useful information about the tasks via a computing device in a different room, the user may not be able to comfortably and / or safely access the computing device. This may be particularly evident in situations where the user is performing a skilled task and / or has otherwise expended energy to be in the user's current situation (e.g., standing on a ladder, working under a vehicle, painting the house, etc.). If the user requests some information while working in such a situation, the user may not be able to properly hear and / or see the rendered content. For example, when the user is working in his garage, he may have a computing device in his garage to view useful content, but when the user is standing at some location in the garage, he may not be able to see the display panel of the computing device. As a result, the user may unfortunately need to pause his work progress to view the display panel to perceive the information provided via the computing device. Furthermore, depending on the amount of time that the content is rendered, a particular user may not have time to adequately perceive the content - given that they must first navigate around various fixtures and / or people to reach the computing device. In such a situation, if the user does not get an opportunity to perceive the content, the user may end up having to re-request the content, thereby draining the computational resources and power of the computing device. Summary of the Invention [Means for solving the problem]
[0003] Implementations described herein relate to a mobile computing device that selectively navigates to a user and adjusts the viewing angle of a display panel of the mobile computing device according to the user's relative location to render certain content to the user. The mobile computing device may include multiple sections (i.e., housing enclosures), and each section may include one or more motors for adjusting the physical position and / or placement of one or more sections. As an example, a user may provide a verbal utterance such as "Assistant, what's on my agenda today?", and in response, the mobile computing device may navigate toward the user's location, adjust the angle of the display panel of the mobile computing device, and render display content that characterizes the user's agenda. However, when a user requests audio content (e.g., audio content without corresponding display content) and the mobile computing device determines that the current location of the mobile computing device corresponds to a distance at which the audio content is audible to the user, the mobile computing device may bypass navigating to the user. In instances when the mobile computing device is not within distance to render audible audio content, the mobile computing device can determine the user's location and navigate toward the user at least until the mobile computing device is within a certain distance to render the audible audio content.
[0004] The mobile computing device can process data based on one or more sensors to determine how to operate motors of the mobile computing device to position portions of the mobile computing device for rendering content. For example, the mobile computing device may include one or more microphones (e.g., an array of microphones) that respond to sounds originating from different directions. One or more processors can process output from the microphones to determine the origin of the sound relative to the mobile computing device. In this manner, if the mobile computing device determines that a user has requested content that should be rendered closer to the user, the mobile computing device can navigate to the user's location as determined based on the output from the microphones. By allowing the mobile computing device to move in this manner, it can provide comfort to users with disabilities who may not be able to efficiently navigate to a computing device for information and / or other media. Furthermore, because the mobile computing device can make decisions regarding when to navigate to a user and when not to, the mobile computing device can conserve power and other computational resources, at least when rendering content. For example, if a mobile computing device navigates to a user indiscriminately with respect to the type of content to be rendered, the mobile computing device may consume more power to navigate to the user as compared to only rendering the content without navigating to the user. Moreover, by rendering content without first navigating to the user when there is no need to do so, unnecessary delays in rendering the content can be avoided.
[0005] In some implementations, a mobile computing device may include an upper housing enclosure that includes a display panel that may have a viewing angle that is adjustable via an upper housing enclosure motor. For example, when the mobile computing device determines that a user is providing oral speech and the viewing angle of the display panel should be adjusted so that the display panel is pointed toward the user, the upper housing enclosure motor may adjust the position of the display panel. When the mobile computing device has completed rendering a particular content, the upper housing enclosure motor may move the display panel back to a resting position, thereby consuming less space than when the display panel was adjusted toward the user.
[0006] In some implementations, the mobile computing device may further include a middle housing enclosure and / or a lower housing enclosure, each capable of housing one or more portions of the mobile computing device. For example, the middle housing enclosure and / or the lower housing enclosure may include one or more cameras for capturing images of the surrounding environment of the mobile computing device. Image data generated based on the output of the camera of the mobile computing device may be used to determine a location of a user to enable the mobile computing device to navigate toward the location when the user provides some commands to the mobile computing device. In some implementations, the middle housing enclosure and / or the lower housing enclosure may be located below the upper housing enclosure when the mobile computing device is operating in a sleep mode, and the middle housing enclosure and / or the lower housing enclosure may include one or more motors for repositioning the mobile computing device, including the upper housing enclosure, when transitioning out of the sleep mode. For example, in response to a user providing an invocation phrase such as "Assistant...", the mobile computing device transitions out of the sleep mode and into an operating mode. During the transition, one or more motors of the mobile computing device can cause the camera and / or the display panel to be pointed toward the user. In some implementations, a first set of motors of the one or more motors can control an orientation of the camera, and a second set of motors of the one or more motors can control a separate orientation of the display panel. In this manner, the camera can have a separate orientation relative to the orientation of the display panel. Additionally or alternatively, the camera can have the same orientation relative to the orientation of the display panel according to the movement of the one or more motors.
[0007] Additionally, during a transition from the compact configuration of the mobile computing device to the extended configuration of the mobile computing device, one or more microphones of the mobile computing device may monitor for further input from the user. When further input is provided by the user during operation of one or more motors of the mobile computing device, noise from the motors may interfere with some frequencies of the verbal input from the user. Thus, to eliminate adverse effects on the quality of the sound captured by the microphones of the mobile computing device, the mobile computing device may modify and / or pause operation of one or more motors of the mobile computing device while the user is providing a subsequent verbal utterance.
[0008] For example, after a user provides an invocation phrase of "Assistant..." and while the motors of the mobile computing device are operating to extend the display panel toward the user, the user may provide a subsequent verbal utterance. The subsequent verbal utterance may be "Show me the security camera in front of the house...," which may cause the automated assistant to invoke an application for viewing live streaming video from the security camera. However, because the motors of the mobile computing device are operating while the user is providing the subsequent verbal utterance, the mobile computing device may determine that the user is providing a verbal utterance and, in response, modify one or more operations of one or more motors of the mobile computing device. For example, the mobile computing device may stop operation of the motors causing the display panel to extend toward the user to remove motor noise that becomes disturbing to the mobile computing device when generating audio data that characterizes the verbal utterance. When the mobile computing device determines that the verbal utterance is no longer being provided by the user and / or is otherwise completed, operation of the one or more motors of the mobile computing device may continue.
[0009] Each motor of the mobile computing device can perform various tasks to achieve some operations of the mobile computing device. In some implementations, the lower housing enclosure of the mobile computing device may include one or more motors connected to one or more wheels (e.g., cylindrical wheels, ball wheels, Mecanum wheels, etc.) for navigating the mobile computing device to a location. The lower housing enclosure may also include one or more other motors for moving a middle housing enclosure of the mobile computing device (e.g., rotating the middle housing enclosure about an axis that is perpendicular to a surface of the lower housing enclosure).
[0010] In some implementations, the mobile computing device can perform one or more different response gestures to indicate that the mobile computing device is receiving input from a user, thereby acknowledging the input. For example, in response to detecting a verbal utterance from a user, the mobile computing device can determine whether the user is within viewing range of a camera of the mobile computing device. If the user is within viewing range of the camera, the mobile computing device can operate one or more motors to invoke a physical movement by the mobile computing device, thereby indicating to the user that the mobile computing device is acknowledging the input.
[0011] The above description is provided as an overview of some implementations of the present disclosure. Further description of those implementations, as well as other implementations, are described in more detail below.
[0012] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform a method such as one or more of the methods described above and / or elsewhere herein. Still other implementations may include one or more computers, including one or more processors, and / or one or more robotic systems operable to execute the stored instructions to perform a method such as one or more of the methods described above and / or elsewhere herein.
[0013] It should be appreciated that all combinations of the above concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Brief description of the drawings]
[0014] [Figure 1A] FIG. 1 illustrates a mobile computing device that navigates autonomously and selectively in response to commands from one or more users. [Figure 1B] FIG. 1 illustrates a mobile computing device that navigates autonomously and selectively in response to commands from one or more users. [Figure 1C] FIG. 1 illustrates a mobile computing device that navigates autonomously and selectively in response to commands from one or more users. [Diagram 2] FIG. 1 illustrates a user providing verbal utterances to a mobile computing device to invoke a response from an automated assistant accessible via the mobile computing device. [Diagram 3]FIG. 13 illustrates a user providing a verbal utterance that causes a mobile computing device to navigate to the user and to place a different housing enclosure of the mobile computing device toward the user to protect the graphical content. [Figure 4] FIG. 1 illustrates a user providing verbal utterances to a mobile computing device that can pause intermittently during navigation to capture any additional verbal utterances from the user. [Diagram 5] FIG. 10 further illustrates a system for operating a computing device that provides access to an automated assistant and autonomously moves toward and / or away from a user in response to spoken utterances. [Figure 6] FIG. 1 illustrates a method for rendering content in a mobile computing device that selectively and autonomously navigates to a user in response to spoken utterances. [Figure 7] FIG. 1 is a block diagram of an exemplary computer system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] 1A, 1B, and 1C show diagrams of a mobile computing device 102 that autonomously and selectively navigates in response to commands from one or more users. Specifically, FIG. 1A shows a perspective view 100 of the mobile computing device 102 in a folded state in which the housing enclosures of the mobile computing device 102 may be closest to one another. In some implementations, the mobile computing device 102 may include one or more housing enclosures, including, but not limited to, one or more of a first housing enclosure 106, a second housing enclosure 108, and / or a third housing enclosure 110. One or more of the housing enclosures may include one or more motors for moving a particular housing enclosure to a particular position and / or toward a particular destination.
[0016] In some implementations, the first housing enclosure 106 may include one or more first motors (i.e., a single motor, or multiple motors) that operate to adjust the orientation of the display panel 104 of the mobile computing device 102. The first motor of the first housing enclosure 106 may be controlled by one or more processors of the mobile computing device 102 and may be powered by a portable power source of the mobile computing device 102. The portable power source may be a rechargeable power source, such as a battery and / or a capacitor, and a power management circuit of the mobile computing device 102 may adjust the output of the portable power source to provide power to the first motor and / or any other motors of the mobile computing device 102. During operation of the mobile computing device 102, the mobile computing device 102 may receive input from a user and determine a response to provide to the user. When the response includes display content, the first motor may adjust the display panel 104 to be directed toward the user. For example, as shown in diagram 120 of FIG. 1B, the display panel 104 may be moved in a direction that increases the angle of separation 124 between the first housing enclosure 106 and the second housing enclosure 108. However, when the response includes audio content without providing corresponding display content, the mobile computing device 102 may remain in a compressed state (as shown in FIG. 1A) without the first motor adjusting the orientation of the display panel 104.
[0017] In some implementations, the second housing enclosure 108 may include one or more second motors for moving the orientation of the first housing enclosure 106 and / or the second housing enclosure 108 relative to the third housing enclosure 110. For example, as shown in diagram 130 of FIG. 1C, the one or more second motors may include a motor that moves the first housing enclosure 106 about an axis 132 that is perpendicular to a surface of the second housing enclosure 108. Thus, the first motor incorporated in the first housing enclosure 106 can modify the angle of separation between the first housing enclosure 106 and the second housing enclosure 108, and the second motor incorporated in the second housing enclosure 108 can rotate the first housing enclosure 106 about the axis 132.
[0018] In some implementations, the mobile computing device 102 may include a third housing enclosure 110 including one or more third motors to move the first housing enclosure 106, the second housing enclosure 108, and / or the third housing enclosure 110 and / or navigate the mobile computing device 102 to one or more different destinations. For example, the third housing enclosure 110 may include one or more motors for moving the second housing enclosure 108 about another axis 134. The another axis 134 may intersect with a rotatable plate 128, which may be attached to a third motor enclosed within the third housing enclosure 110. Additionally, the rotatable plate 128 may be connected to an arm 126, on which the second housing enclosure 108 may be mounted. In some implementations, a second motor enclosed within the second housing enclosure 108 can be connected to an arm 126, which can act as a fulcrum to allow an angle of separation 136 between the third housing enclosure 110 and the second housing enclosure 108 to be adjusted.
[0019] In some implementations, the mobile computing device 102 may include another arm 126 that can act as another fulcrum or other device to assist the second motor by adjusting the angle of separation 124 between the first housing enclosure 106 and the second housing enclosure 108. The placement of the mobile computing device 102 may depend on the environment in which the mobile computing device 102 and / or another computing device receives input. For example, in some implementations, the mobile computing device 102 may include one or more microphones that are oriented in different directions. For example, the mobile computing device 102 may include a first set of one or more microphones 114 oriented in a first direction and a second set of one or more microphones 116 oriented in a second direction different from the first direction. The microphones may be attached to the second housing enclosure 108 and / or any other housing enclosure or combination of housing enclosures.
[0020] When the mobile computing device 102 receives input from a user, signals from the one or more microphones may be processed to determine a location of the user relative to the mobile computing device 102. When the location is determined, one or more processors of the mobile computing device 102 may cause one or more motors of the mobile computing device 102 to position the housing enclosure such that the second set of microphones 116 is pointed toward the user. Additionally, in some implementations, the mobile computing device 102 may include one or more cameras 112. The cameras 112 may be mounted in the second housing enclosure and / or any other housing enclosure or combination of housing enclosures. In response to receiving the input and determining the location of the user by the mobile computing device 102, the one or more processors may cause one or more motors to position the mobile computing device 102 such that the cameras 112 are pointed toward the user.
[0021] As one non-limiting example, the mobile computing device 102 may determine that the child has directed an oral utterance at the mobile computing device 102 while the mobile computing device is on a table that may be higher than the child. In response to receiving the oral utterance, the mobile computing device 102 may use signals from one or more of the microphone 114, microphone 116, and / or camera 112 to determine the location of the child relative to the mobile computing device 102. In some implementations, processing of the audio and / or video signals may be offloaded to a remote device, such as a remote server, via a network to which the mobile computing device 102 is connected. The mobile computing device 102 may determine, based on its processing, that an anatomical feature (e.g., an eye, an ear, a face, a mouth, and / or any other anatomical feature) is located below the table. Thus, to direct the display panel 104 toward the user, one or more motors of the mobile computing device 102 may position the display panel 104 in a direction that is below the table.
[0022] 2 is a diagram 200 in which a user 202 provides verbal utterances to a mobile computing device 204 to invoke a response from an automated assistant accessible via the mobile computing device 204. Specifically, the user 202 can provide verbal utterances 210, which can be captured via one or more microphones of the mobile computing device 204. The mobile computing device 204, and / or a computing device in communication with the mobile computing device 204, can process audio data characterizing the verbal utterances 210. Based on that processing, the mobile computing device 204 can determine one or more actions that have been requested by the user 202. Additionally, the mobile computing device 204 can determine whether the performance of the one or more actions involves rendering graphical content on a display panel of the mobile computing device 204. When one or more actions do not involve rendering graphical content for the user 202, the mobile computing device 204, in some implementations, can bypass navigating around obstacles, such as a couch 208, within the room 206 in which the user 202 is located.
[0023] Alternatively, the mobile computing device 204 may determine that one or more actions involve rendering audio content and further determine whether the user 202 is located within distance to effectively render the audible content for the user 202. In other words, the mobile computing device 204 may determine whether the user 202 is located close enough to the mobile computing device 204 to hear any audio output generated at the mobile computing device 204. If the user 202 is within distance to audibly render the audio content, the mobile computing device 204 may bypass navigating closer to the user 202. However, if the user is not within distance to audibly render the audio content, the mobile computing device 204 may control one or more motors of the mobile computing device 204 to navigate closer to the location of the user 202.
[0024] When the mobile computing device 204 reaches a distance to audibly render audio content for the user 202 or otherwise determines that the mobile computing device 204 is already within a distance to audibly render audio content, the mobile computing device 204 can provide a response output 212. For example, when the spoken utterance includes natural language content such as, "Assistant, what's on my schedule for today?", the automated assistant can provide a response output such as, "Okay, you're going to meet Renee for coffee at 9:00 AM and you're going to play games with Sol and Darren at 11:30 AM." In this manner, power can be conserved in the mobile computing device 204 by selectively navigating or not navigating to the user 202 depending on one or more actions requested by the user 202. Additionally, the ability of the mobile computing device 204 to navigate to the user 202 and provide a response when the user 202 is unable to reach the mobile computing device 204 can eliminate the need for the user 202 to stop what they are doing in some situations.
[0025] 3 shows a diagram 300 in which a user 302 provides a verbal utterance 310 that causes a mobile computing device 304 to navigate to the user 302 and to place a different housing enclosure of the mobile computing device 304 towards the user 302 to protect the graphical content. The mobile computing device 304 can selectively navigate to the user 302 according to whether the user 302 is requesting that a particular type of content be provided to the user 302. For example, as provided in FIG. 2, the mobile computing device 204 can bypass navigating to the user 202 when the user is requesting audio content and the mobile computing device 204 determines that the user 202 is within a threshold distance to audibly render the audio content for the user 202. However, when the user requests graphical content, such as in FIG. 3, the mobile computing device 304 can navigate to the user 302 to present the graphical content on a display panel of the mobile computing device 304.
[0026] For example, as provided in FIG. 3, the user 302 may provide an oral utterance 310 such as, "Assistant, show me yesterday's security video." In response to receiving the oral utterance 310, the mobile computing device 304 may generate audio data characterizing the oral utterance, process the audio data at the mobile computing device 304, and / or transmit the audio data to another computing device for processing. Based on the processing of the audio data, the mobile computing device 304 may determine that the user 302 is requesting that the mobile computing device 304, and / or any other display-enabled device, provide playback of a security video for the user 302. Based on this determination, the mobile computing device 304 may determine the location of the user 302 relative to the mobile computing device 304. Additionally, or in the alternative, the mobile computing device 304 may identify one or more obstacles present in a room 306 through which the mobile computing device 304 successfully navigates to reach the user 302. For example, by using image data captured via a camera of the mobile computing device 304, it may be determined that a couch 308 separates the user 302 from the mobile computing device 304. Using this image data, the mobile computing device 304, and / or a remote computing device that processes the image data, may generate a route 312 for reaching the user 302.
[0027] In some implementations, the mobile computing device 304 may be connected to a local network to which other computing devices are connected. When the mobile computing device 304 determines that it is unable to navigate to the user 302 due to one or more obstacles and / or is unable to navigate to the user within a threshold amount of time, an automated assistant accessible via the mobile computing device 304 may identify other display-capable devices connected on the local area network. For example, the automated assistant may determine that the television 314 is located in the same room as the user 302 and that the mobile computing device 304 is unable to reach the user. Based on these determinations, the automated assistant may cause the requested display content to be rendered on the television 314 and / or any other computing devices located in the room 306 that are display-capable.
[0028] However, when the mobile computing device 304 can navigate to the route 312 to reach the user 302, the mobile computing device 304 can identify one or more anatomical features of the user 302 when it reaches the user 302 and / or when the mobile computing device 304 is en route to the location of the user 302. For example, when the mobile computing device 304 is navigating the route 312 to reach the user 302, the mobile computing device 304 can determine that the user 302 is within a viewing window of the camera of the mobile computing device 304. In response to this determination, the mobile computing device 304 can use image data captured via the camera to identify the eyes, mouth, and / or ears of the user 302. Based on the identification of one or more of these anatomical features of the user 302, the mobile computing device 304 can cause one or more motors of the mobile computing device 304 to move a display panel of the mobile computing device 304 toward the direction of the user 302.
[0029] For example, when a display panel is connected to the first housing enclosure of the mobile computing device 304, one or more first motors of the mobile computing device 304 can cause an angle of separation between the first housing enclosure and the second housing enclosure to increase. Additionally, based on the determined location of the anatomical features of the user 302 relative to the mobile computing device 304, one or more motors of the mobile computing device 304 can further raise the height of the mobile computing device 304 such that the display panel is more easily viewable by the user 302. For example, one or more second motors of the mobile computing device 304 can cause the second housing enclosure of the mobile computing device 304 to have an increased angle of separation relative to the third housing enclosure of the mobile computing device 304. These increases in the angle of separation can cause the mobile computing device 304 to transform from a compressed state to an expanded state, thereby raising the height of the mobile computing device 304. When the mobile computing device 304 has completed rendering the graphical content in accordance with the spoken utterance 310, the mobile computing device 304 may return to the folded state to conserve stored energy in the mobile computing device's 304 rechargeable power source.
[0030] FIG. 4 illustrates a diagram 400 in which a user 402 provides verbal utterances to a mobile computing device 404, which may pause intermittently during navigation to capture any additional verbal utterances from the user 402. The mobile computing device 404 may provide access to an automated assistant, which may respond to a variety of different inputs from one or more users. The user 402 may provide verbal utterances, which may be processed at the mobile computing device 404 and / or another computing device associated with the mobile computing device 404. The mobile computing device 404 may include one or more microphones, which may provide an output signal in response to the verbal input from the user. To filter out noise that may otherwise affect the verbal input, the mobile computing device 404 may determine whether the user 402 is providing verbal input when one or more motors of the mobile computing device 404 are operating. In response to determining that oral speech is being provided while one or more motors are operating, the mobile computing device 404 may cause the one or more motors to enter a low power state to reduce the amount of noise generated by the one or more motors.
[0031] For example, the user 402 may provide a spoken utterance 410, such as, "Assistant, make a video call to my brother...", and the mobile computing device 404 may receive the spoken utterance 410 and determine that the spoken utterance 410 corresponds to an action involving a camera and / or rendering of graphical content. In response to determining that the action involves the mobile computing device 404's camera and rendering of graphical content, the mobile computing device 404 may navigate toward the user's 402 location. As the mobile computing device 404 traverses the first portion 412 of the route, the user 402 may provide a subsequent spoken utterance 410, such as, "...and also, set the house alarm on." While traversing the first portion 412 of the route, the mobile computing device 404 may determine that the user 402 is providing a subsequent spoken utterance. In response, the mobile computing device 404 may cause one or more motors of the mobile computing device 404 to enter a lower power state compared to the power state in which the one or more motors were operating when the mobile computing device 404 was traversing the first portion 412 of the route. For example, the one or more motors may pause their respective operation, thereby causing the mobile computing device 404 to come to rest after the first portion 412 of the route for the user 402.
[0032] The mobile computing device 404 and / or the automated assistant may later determine that the verbal utterance is complete and / or is otherwise no longer directed at the mobile computing device 404, and the mobile computing device 404 may proceed to traverse a second portion 414 of the route toward the location of the user 402. The second portion 414 of the route may include navigating through a room 406, including a couch 408 and / or other obstacles. The mobile computing device 404 may use a camera to identify such obstacles, as well as the user 402, along with prior permission from the user. In some implementations, while the mobile computing device 404 is traversing the second portion 414 of the route, the automated assistant may initiate the performance of other actions requested by the user 402. Specifically, while the mobile computing device 404 is traversing the second portion 414 of the route toward the location of the user 402, the mobile computing device 404 may initiate the setting of a house alarm. This action may be initiated based on a determination that the action does not involve rendering graphical content that the user would want to see and / or does not involve rendering audio data that the mobile computing device 404 has attempted to render at a distance that would allow the audio content to be audible to the user 402.
[0033] 5 illustrates a system 500 for operating a computing device 518 to selectively navigate to a user to render content to the user and to toggle motor operation according to whether the user is providing verbal utterances to the computing device 518. The automated assistant 504 can operate as part of an assistant application provided on one or more computing devices, such as the computing device 518 and / or the server device 502. The user can interact with the automated assistant 504 through an assistant interface, which can be a microphone, a camera, a touch screen display, a user interface, and / or any other device capable of providing an interface between a user and an application.
[0034] For example, a user can initialize the automated assistant 504 by providing verbal, textual, and / or graphical input to the assistant interface, causing the automated assistant 504 to perform functions (e.g., provide data, control peripheral devices, access agents, generate input and / or output, etc.). The computing device 518 can include a display device, which can be a display panel including a touch interface for receiving touch input and / or gestures to enable a user to control applications of the computing device 518 via the touch interface. In some implementations, the computing device 518 can lack a display device, but can thereby provide audible user interface output without providing graphical user interface output. Additionally, the computing device 518 can provide a user interface, such as a microphone, to receive verbal natural language input from a user. In some implementations, the computing device 518 can include a touch interface and can lack a camera, but can in some cases include one or more other sensors.
[0035] The computing device 518 and / or other computing devices may be in communication with the server device 502 over a network 536, such as the Internet. Additionally, the computing device 518 and other computing devices may be in communication with each other over a local area network (LAN), such as a WiFi network. The computing device 518 may offload computational tasks to the server device 502 to conserve computational resources at the computing device 518. For example, the server device 502 may host an automated assistant 504, and the computing device 518 may send inputs received at one or more assistant interfaces 520 to the server device 502. However, in some implementations, the automated assistant 504 may be hosted at the computing device 518 as a client automated assistant 522.
[0036] In various implementations, all or less than all aspects of the automated assistant 504 may be implemented on the computing device 518. In some of those implementations, aspects of the automated assistant 504 are implemented via a client automated assistant 522 on the computing device 518, which interfaces with a server device 502, which implements other aspects of the automated assistant 504. The server device 502 may serve multiple users and their associated assistant applications, possibly via multiple threads. In implementations in which all or less than all aspects of the automated assistant 504 are implemented via a client automated assistant 522 on the computing device 518, the client automated assistant 522 may be an application that is separate from the operating system of the computing device 518 (e.g., installed "on top of" the operating system) - or alternatively, may be implemented directly by the operating system of the computing device 518 (e.g., an operating system application, but considered integral to the operating system).
[0037] In some implementations, the automated assistant 504 and / or the client automated assistant 522 may include an input processing engine 506, which may employ multiple different modules to process input and / or output for the computing device 518 and / or the server device 502. For example, the input processing engine 506 may include a voice processing module 508, which may process audio data received at the assistant interface 520 to identify text embedded in the audio data. The audio data may be transmitted from the computing device 518 to the server device 502, for example, to conserve computational resources at the computing device 518.
[0038] The process for converting the audio data to text may include a speech recognition algorithm, which may employ neural networks and / or statistical models to identify groups of audio data that correspond to words or phrases. The text converted from the audio data may be parsed by a data parsing module 510 and made available to the automated assistant as text data that may be used to generate and / or identify command phrases from the user. In some implementations, the output data provided by the data parsing module 510 may be provided to a parameter module 512 to determine whether the user has provided input corresponding to a particular action and / or routine that may be performed by the automated assistant 504 and / or an application or agent that may be accessed by the automated assistant 504. For example, the assistant data 516 may be stored as client data 538 at the server device 502 and / or the computing device 518 and may include data defining one or more actions that may be performed by the automated assistant 504 and / or the client automated assistant 522, as well as parameters necessary to perform those actions.
[0039] 5 further illustrates a system 500 for operating a computing device 518 that provides access to an automated assistant 504 and autonomously moves toward and / or away from a user in response to verbal utterances. The computing device 518 may be powered by one or more power sources 526, which may be rechargeable and / or may enable the computing device 518 to be portable. A motor control engine 532 may be powered by the power source 526 and may determine when to control one or more motors of the computing device 518. For example, the motor control engine 532 may determine one or more operational statuses of the computing device 518 and control one or more motors of the computing device 518 to reflect the one or more operational statuses. For example, when the computing device 518 receives verbal speech at the assistant interface 520 of the computing device 518, the motor control engine 532 can determine that a user is providing verbal speech and cause one or more motors to operate to facilitate indicating that the computing device 518 is acknowledging the verbal speech. The one or more motors can be caused to vibrate and / or dance, for example, when the computing device 518 receives the verbal speech. Alternatively or additionally, the motor control engine 532 can cause the one or more motors to move the computing device 518 back and forth via a ground wheel of the computing device 518 to indicate that the computing device 518 is downloading and / or uploading data over a network.Alternatively or additionally, the motor control engine 532 can cause one or more motors to position the housing enclosure of the computing device 518 into a compressed or relaxed state to indicate that the computing device 518 is operating in a low power mode and / or a sleep mode.
[0040] When operating in a sleep mode, the computing device 518 may monitor for an invocation phrase spoken by a user and / or may perform voice activity detection. When the computing device is performing voice activity detection, the computing device 518 may determine whether an input to a microphone corresponds to a human. Furthermore, voice activity detection may be performed when the computing device 518 is controlled by one or more motors of the computing device 518. In some implementations, the threshold for determining whether a human voice is detected may include a threshold for when the one or more motors are operating and another threshold for when the one or more motors are not operating. For example, when the computing device 518 is in a sleep mode, voice activity detection may be performed according to a first threshold, and the first threshold may be met when a first percentage of the incoming noise corresponds to a human voice. However, when the computing device 518 is in an awake mode, voice activity detection may be performed according to a second threshold, which may be met when a second percentage of the incoming noise corresponds to a human voice, the second percentage of the incoming noise being higher than the first percentage of the incoming noise. Furthermore, when the computing device 518 is in a wake mode and one or more motors are operating to reposition the computing device 518 and / or navigate the computing device 518, voice activity detection may be performed according to a third threshold, which may be met when a third percentage of the incoming noise corresponds to a human voice. The third percentage of the incoming noise may be greater than and / or equal to the second percentage of the incoming noise and / or the first percentage of the incoming noise.
[0041] In some implementations, in response to a determination by the computing device 518 that a human voice has been detected, the spatial processing engine 524 can process incoming data from one or more sensors to determine where the human voice is coming from. Spatial data characterizing the location of a source of the human voice, such as a user, can be generated by the spatial processing engine 524 and communicated to the location engine 530. The location engine 530 can use the spatial data to generate a route for navigating the computing device 518 from its current location at the computing device 518 to the location of the source of the human voice. The route data can be generated by the location engine 530 and communicated to the motor control engine 532. The motor control engine 532 can use the route data to control one or more motors of the computing device 518 to navigate to the location of the user and / or the source of the human voice.
[0042] In some implementations, the spatial processing engine 524 can process incoming data from one or more sensors of the computing device 518 to determine whether the computing device 518 is located within a distance from a source of a human voice to render audible audio. For example, the automated assistant can receive a request from a user and can determine one or more actions requested by the user. The one or more actions can be communicated to the content rendering engine 534, which can determine whether the user is requesting that audio content be rendered, that graphical content be rendered, and / or that either audio content or graphical content be rendered. In response to determining that the user has requested that audio content be rendered, the spatial processing engine 524 can determine whether the computing device 518 is located within a distance from the user to generate audio content that is audible to the user. When the computing device 518 determines that the computing device 518 is not within a distance to generate audible content, the computing device 518 can control one or more motors to navigate the computing device 518 to be within a distance to generate audible content.
[0043] Alternatively, in response to determining that the user has requested that graphical content be rendered, the spatial processing engine 524 may determine whether the computing device 518 is located within another distance from the user to generate graphical content that is visible to the user. When the computing device 518 determines that the computing device 518 is not within another distance to generate visible graphical content, the computing device 518 may control one or more motors to navigate the computing device 518 to be within another distance to generate visible graphical content. In some implementations, the amount of distance between the user and the computing device 518 may be based on a particular property of the graphical content when the computing device 518 is rendering the graphical content. For example, when the graphical content includes text that is X size, the computing device 518 may navigate to be within m distance from the user. However, when the graphical content includes text that is Y size that is less than X size, the computing device 518 may navigate to be within N distance from the user that is less than M. Alternatively or additionally, the distance between the user and the computing device 518 may be based on the type of content to be generated. For example, the computing device 518 may navigate up to a distance H from the user when the graphical content to be rendered includes video content, and up to a distance K from the user when the graphical content to be rendered includes still images, where H is less than K.
[0044] FIG. 6 illustrates a method 600 for rendering content at a mobile computing device that selectively and autonomously navigates to a user in response to a spoken utterance. The method 600 may be implemented by one or more computing devices, applications, and / or any other devices or modules capable of responding to the spoken utterance. The method 600 may include an operation 602 of determining whether a spoken utterance has been received from a user. The mobile computing device may include one or more microphones with which the mobile computing device can detect verbal input from the user. Additionally, the mobile computing device may provide access to an automated assistant, which may initiate actions and / or render content in response to the provision of one or more inputs by the user. For example, a user may provide a verbal utterance such as "Assistant, send a video message to Megan." The mobile computing device may generate audio data based on the verbal utterance and cause the audio data to be processed to identify one or more actions requested by the user (e.g., initiating a video call to a contact).
[0045] When no verbal speech is detected at the mobile computing device, one or more microphones of the mobile computing device may be monitored for verbal input. However, when verbal speech is received, method 600 may proceed from operation 602 to operation 604. Operation 604 may include determining whether the requested action involves rendering graphical content. The graphical content may be, but is not limited to, media provided by an application, streaming data, video recorded by a camera accessible to the user, and / or any other video data that may or may not be associated with corresponding audio data. For example, when a user requests that a video message be provided to another person, the mobile computing device may determine that the requested action involves rendering graphical content because generating the video message may involve rendering a video preview of the video message and rendering a video stream of the recipient (e.g., "Megan").
[0046] When it is determined that the requested action involves rendering graphical content, the method 600 may proceed from operation 604 to operation 608. Operation 608 may include determining whether the user is within or at a distance to perceive the graphical content. That is, the operation may determine whether the user's location relative to the mobile computing device satisfies a distance condition. The distance condition may be predetermined and may be fixed, for example, for all graphical content. Alternatively, the distance condition may vary depending on the particular graphical content (i.e., the distance condition may be determined based on the graphical content). For example, a display of basic content, which may be displayed in a large font, may be associated with a different distance condition compared to a display of detailed or densely presented content. In other words, the mobile computing device, and / or a server in communication with the mobile computing device, may process data to determine whether the user is able to perceive the graphical content to be displayed on the mobile computing device. For example, the mobile computing device may include a camera that captures image data, and the image data may characterize the user's location relative to the mobile computing device. The mobile computing device can use the image data to determine the proximity of the user to the mobile computing device, thereby determining whether the user can comfortably view the display panel of the mobile computing device. When the mobile computing device determines that the user is not within distance to perceive the graphical content (i.e., a distance condition associated with the content is met), method 600 can proceed to operation 610.
[0047] Operation 610 may include causing the mobile computing device to move within a distance for perceiving the graphical content. In other words, the mobile computing device may operate one or more motors to navigate the mobile computing device toward the user at least until the mobile computing device reaches or is within a distance for the user to perceive the graphical content. When it is determined that the user is within a distance for perceiving (e.g., being able to see and / or read) the graphical content, method 600 may proceed from operation 608 to operation 612.
[0048] When it is determined that the requested action does not involve rendering graphical content, the method 600 may proceed from operation 604 to operation 606. Operation 606 may include determining whether the requested action involves rendering audio content. Audio content may include any output from the mobile computing device and / or any other computing device that may be audible to one or more users. When it is determined that the requested action involves rendering audio content, the method 600 may proceed from the operation at 606 to operation 616. Otherwise, when it is determined that the requested action does not involve rendering audio content and / or graphical content, the method 600 may proceed to operation 614, where one or more requested actions are initialized in response to the verbal utterance.
[0049] Operation 616 may include determining whether the user is within distance to perceive the audio content. In other words, the mobile computing device may determine whether the user's current location will allow the user to hear audio generated at the mobile computing device or another computing device that can render the audio content in response to the oral utterance. For example, the mobile computing device may generate audio data and / or image data from which the user's location relative to the mobile computing device may be estimated. When the user's estimated distance from the mobile computing device is not within distance to perceive the audio content, method 600 may proceed from operation 616 to operation 618.
[0050] Operation 618 may include causing the mobile computing device to move within distance to perceive the audio content. Alternatively or additionally, operation 618 may include determining whether one or more other computing devices are within distance from the user to render the audio content. Thus, if another computing device is located within distance to render audible audio content for the user, the determination at operation 616 may be affirmatively satisfied and method 600 may proceed to operation 620. Otherwise, the mobile computing device may move closer to the user such that the mobile computing device is within distance to perceive the audio content generated by the mobile computing device. When the mobile computing device is within distance to receive the audio content, method 600 may proceed from operation 616 to operation 620.
[0051] In instances where the requested action involves graphical content and the mobile computing device has moved within a distance for the user to perceive the graphical content, the method 600 may proceed from operation 608 to operation 612. Operation 612 may include causing the mobile computing device to move a display panel to be pointed toward the user. The display panel may be controlled by one or more motors attached to one or more housing enclosures of the mobile computing device. For example, the one or more motors may be attached to a first housing enclosure and may operate to adjust the angle of the display panel. With permission from the user, image data and / or audio data captured at the mobile computing device and / or any other computing device may be processed to identify one or more anatomical features of the user, such as the user's eyes. Based on the identification of the anatomical features, the one or more motors controlling the angle of the display panel may be operated to move the display panel such that the display panel projects the graphical content toward the anatomical features of the user. In some implementations, one or more other motors of the mobile computing device may further adjust the height of the display panel of the mobile computing device. Thus, one or more motors and / or one or more other motors can operate simultaneously to move the display panel so that it is within the user's field of view and / or is oriented toward the user's anatomical features.
[0052] When the mobile computing device completes moving the display panel to face the user, the method 600 may proceed from operation 612 to operation 620. Operation 620 may include causing the requested content to be rendered and / or causing the requested action to be performed. For example, when a user provides a verbal utterance requesting an automated assistant to turn on the lights in a house, this action may involve controlling an IoT device without rendering audio and / or display content, thereby allowing the mobile computing device to bypass moving toward the user's direction. However, when the verbal utterance includes a request for an audio stream and / or a video stream to be provided via the mobile computing device, the mobile computing device may move toward the user and / or verify that the user is within distance to perceive the content. The mobile computing device may then render the content for the user. In this manner, delays that may otherwise be caused by having the user first request the user to navigate to the user before the mobile computing device can render the content. Furthermore, the mobile computing device can conserve computational resources by choosing whether or not to navigate to a user depending on the type of content to be rendered for the user. Such computational resources, such as power and processing bandwidth, may otherwise be wasted if the mobile computing device indiscriminately navigated to a user without regard for the action being requested.
[0053] 7 is a block diagram of an exemplary computer system 710. The computer system 710 typically includes at least one processor 714 that communicates with several peripheral devices via a bus subsystem 712. These peripheral devices may include, for example, a storage subsystem 724 including a memory 725 and a file storage subsystem 726, a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices enable user interaction with the computer system 710. The network interface subsystem 716 provides an interface to external networks and is coupled to corresponding interface devices in other computer systems.
[0054] The user interface input devices 722 may include keyboards and pointing devices such as a mouse, trackball, touchpad, or graphics tablet, scanners, touch screens integrated into displays, and audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 710 or onto a communications network.
[0055] The user interface output devices 720 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from computer system 710 to a user or to another machine or computer system.
[0056] Storage subsystem 724 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 724 may include logic for performing selected aspects of method 600 and / or implementing one or more of system 500, mobile computing device 102, mobile computing device 204, mobile computing device 304, mobile computing device 404, automated assistant, computing device 518, server device 502, and / or any other applications, devices, apparatus, and / or modules described herein.
[0057] These software modules are generally executed by the processor 714 alone or in combination with other processors. The memory 725 used within the storage subsystem 724 may include several memories, including a main random access memory (RAM) 730 for storing instructions and data during program execution, and a read-only memory (ROM) 732 in which fixed instructions are stored. The file storage subsystem 726 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of some implementations may be stored by the file storage subsystem 726 within the storage subsystem 724 or within other machines accessible by the processor 714.
[0058] Bus subsystem 712 provides a mechanism for allowing the various components and subsystems of computer system 710 to communicate with each other as intended. Although bus subsystem 712 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0059] The computer system 710 can be of different types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computer system 710 shown in Figure 7 is only a specific example to illustrate some implementations. Many other configurations of the computer system 710 are possible, having more or fewer components than the computer system shown in Figure 7.
[0060] In situations where the systems described herein may collect or utilize personal information about users (or as they are often referred to herein as “participants”), the user may be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or to control whether and / or how to receive content from a content server that may be more relevant to the user. Also, certain data may be treated in one or more ways before being stored or used such that personally identifiable information is removed. For example, the user's identity may be treated such that personally identifiable information cannot be determined about the user, or the user's geographic location may be generalized, in which case geographic location information is obtained (such as to the city, zip code, or state level) such that the user's specific geographic location cannot be determined. Thus, the user may have control over how information is collected and / or used about the user.
[0061] Although several implementations have been described and illustrated herein, various other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein may be utilized, and each such variation and / or modification is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the particular application or applications for which the teachings are used.
[0062] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, numerous equivalents to the specific implementations described herein. Accordingly, it should be understood that the above implementations are presented by way of example only, and that within the scope of the appended claims and their equivalents, implementations may be practiced other than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure, provided such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0063] In some implementations, the method is described as including operations such as determining that a user has provided an oral utterance based on an input to one or more microphones of the mobile computing device, the mobile computing device including one or more first motors that move the mobile computing device through an area. The method may further include determining that the user is requesting the mobile computing device to perform an action associated with the automated assistant rendering content via one or more speakers and / or a display panel of the mobile computing device based on the input to the one or more microphones. The method may further include determining a location of the user relative to the mobile computing device based on the input to the one or more microphones, additional input to the one or more microphones, and / or one or more other sensors of the mobile computing device. The method may further include causing a first motor of the mobile computing device to move the mobile computing device toward the user's location and causing the display panel to render the graphical content to facilitate performing the action when the content requested by the user to be rendered at the mobile computing device includes graphical content and when the determined location meets a particular distance condition.
[0064] In some implementations, the method may further include, when the content requested by the user to be rendered at the mobile computing device includes audio content, determining whether the mobile computing device is within a distance from the user to audibly render the audio content for the user. When the mobile computing device is not within a distance from the user to audibly render the audio content for the user, the method may further include, when the mobile computing device is not within a distance from the user to audibly render the audio content for the user, based on a determination that the mobile computing device is not within a distance from the user, causing one or more first motors of the mobile computing device to move the mobile computing device toward a location of the user and causing one or more speakers of the mobile computing device to render the audio content to facilitate performing the action.
[0065] In some implementations, the method may further include, when the content requested by the user to be rendered on the mobile computing device includes graphical content and when the determined location satisfies the distance condition, causing one or more second motors of the mobile computing device to move a display panel of the mobile computing device toward the user to facilitate rendering the graphical content. In some implementations, the method may further include determining whether a user is providing a subsequent oral utterance when the one or more first motors and / or the one or more second motors of the mobile computing device are operating, and causing the one or more first motors and / or the one or more second motors to transition to a reduced power state when the subsequent oral utterance is received while the one or more first motors and / or the one or more second motors of the mobile computing device are operating, the reduced power state corresponding to a state in which the one or more first motors and / or the one or more second motors consume less power than another and / or a previous state of the one or more first motors and / or the one or more second motors.
[0066] In some implementations, the method may further include causing the one or more first motors and / or the one or more second motors of the mobile computing device to transition from the reduced power state to another operating state to facilitate moving the display panel and / or moving the mobile computing device toward a location of the user when subsequent oral utterances are no longer received while the one or more first motors and / or the one or more second motors of the mobile computing device are operating. In some implementations, the method may further include identifying, in response to receiving input to the microphone and using a camera of the mobile computing device, an anatomical feature of the user, where causing the one or more second motors to move the display panel includes causing the display panel to be directed toward the anatomical feature of the user. In some implementations, the method may further include, when the content requested by the user corresponds to the graphical and / or audio content to be rendered on the mobile computing device, causing one or more third motors of the mobile computing device to move a camera of the mobile computing device to be pointed toward the user based on the correspondence of the content requested by the user to the graphical and / or audio content.
[0067] In some implementations, the display panel is mounted to a first housing enclosure of the mobile computing device, the camera is mounted to a second housing enclosure of the mobile computing device, and the one or more third motors are at least partially enclosed within the third housing enclosure of the mobile computing device. In some implementations, causing the one or more third motors of the mobile computing device to move the camera of the mobile computing device in the direction of the user includes causing the second housing enclosure of the mobile computing device to rotate about an axis that intersects the third housing enclosure of the mobile computing device. In some implementations, a fourth motor is at least partially enclosed in the second housing enclosure and controls radial movement of the second housing enclosure relative to the third housing enclosure, and the method further includes causing the fourth motor to effect radial movement of the second housing enclosure such that the second housing enclosure changes an angle of separation from the third housing enclosure when content requested by the user corresponds to graphical content and / or audio content to be rendered on the mobile computing device.
[0068] In some implementations, the method may further include, in response to receiving an input to the microphone and using a camera of the mobile computing device, identifying an anatomical feature of the user when content requested by the user to be rendered on the mobile computing device corresponds to graphical and / or audio content, and determining an angle of separation of the second housing enclosure relative to the third housing enclosure based on the identification of the anatomical feature of the user, the angle of separation corresponding to an angle at which the camera is directed to the anatomical feature of the user. In some implementations, a fifth motor is at least partially enclosed in the first housing enclosure and controls another radial movement of the first housing enclosure relative to the second housing enclosure, and the method further includes, when content requested by the user to be rendered on the mobile computing device corresponds to graphical and / or audio content, causing the fifth motor to effect another radial movement of the first housing enclosure such that the first housing enclosure reaches another angle of separation with the second housing enclosure.
[0069] In some implementations, the method may further include, when the content requested by the user to be rendered at the mobile computing device corresponds to graphical and / or audio content, identifying an anatomical feature of the user in response to receiving input to the microphone and using a camera of the mobile computing device, and determining, based on the identification of the anatomical feature of the user, a different angle of separation of the first housing enclosure relative to the second housing enclosure, the different angle of separation corresponding to a different angle at which the display panel is oriented toward the anatomical feature of the user. In some implementations, determining a location of the user relative to the mobile computing device includes determining, using output from a plurality of microphones of the mobile computing device, that the location includes a plurality of different people, and determining, using other output from the camera of the mobile computing device, that the user is one of the people of the plurality of different people. In some implementations, the method may further include, after rendering the display content and / or audio content, causing one or more second motors to reduce a height of the mobile computing device by moving a first housing enclosure and a display panel of the mobile computing device toward a second housing enclosure of the mobile computing device.
[0070] In other implementations, the method is described as including operations such as determining that a user has provided verbal utterances to the mobile computing device based on input to one or more microphones of the mobile computing device. The method may further include causing one or more motors of the mobile computing device to move a display panel attached to a first housing enclosure of the mobile computing device away from a second housing enclosure of the mobile computing device in response to providing the verbal utterances to the mobile computing device. The method may further include determining whether another verbal utterance is directed to the mobile computing device while the one or more motors are moving the first housing enclosure away from the second housing enclosure. When it is determined that the other verbal utterance is directed to the mobile computing device, the method may further include causing the one or more motors to transition to a low power state while the other verbal utterance is directed to the mobile computing device, and causing an automated assistant accessible via the mobile computing device to initialize performance of an action based on the other verbal utterance. The method may further include causing the one or more motors to complete moving the first enclosure away from the second housing enclosure when the other oral utterance is completed and / or is no longer directed at the mobile computing device.
[0071] In some implementations, the method may further include, in response to providing the oral utterance to the mobile computing device, causing one or more second motors of the mobile computing device to drive the mobile computing device toward the location of the user. In some implementations, the method may further include, when another oral utterance is determined to be directed at the mobile computing device, causing the one or more second motors of the mobile computing device to pause driving the mobile computing device toward the location of the user, and when the another oral utterance is completed and / or is no longer directed at the mobile computing device, causing the one or more second motors of the mobile computing device to continue driving the mobile computing device toward the location of the user.
[0072] In some implementations, the method may further include, when the one or more second motors have completed driving the mobile computing device toward the user's location, causing, based on the verbal utterance and / or another verbal utterance, one or more third motors of the mobile computing device to move the second housing enclosure away from the third housing enclosure of the mobile computing device and to move a camera of the mobile computing device toward the user. In some implementations, the method may further include, when the one or more second motors have completed driving the mobile computing device toward the user's location, causing, based on the verbal utterance and / or another verbal utterance, one or more fourth motors to rotate the first housing enclosure about an axis that intersects a surface of the third housing enclosure to facilitate orienting the display panel toward the user.
[0073] In yet other implementations, the method is described as including operations such as determining that a user has provided an oral utterance to the mobile computing device based on input to one or more microphones of the mobile computing device. The method may further include causing one or more motors of the mobile computing device to move the mobile computing device toward the user's location in response to providing the oral utterance to the mobile computing device. The method may further include determining whether another oral utterance is directed to the mobile computing device while the one or more motors are moving the mobile computing device toward the user's location. When it is determined that the other oral utterance is directed to the mobile computing device, the method may further include causing the one or more motors to transition to a low power state while the other oral utterance is directed to the mobile computing device, and causing an automated assistant accessible via the mobile computing device to initialize performance of an action based on the other oral utterance. The method may further include causing the one or more motors to continue moving the mobile computing device toward the user's location when the other oral utterance is completed and / or is no longer directed at the mobile computing device.
[0074] In some implementations, the method may further include, in response to providing the verbal utterance to the mobile computing device, causing one or more second motors of the mobile computing device to move a display panel attached to a first housing enclosure of the mobile computing device away from a second housing enclosure of the mobile computing device. In some implementations, the method may further include, when it is determined that another verbal utterance is directed to the mobile computing device, causing one or more second motors to transition to a low power state while the another verbal utterance is directed to the mobile computing device, and causing an automated assistant accessible via the mobile computing device to initialize implementation of an action based on the another verbal utterance. In some implementations, the method may further include, when the another verbal utterance is completed and / or is no longer directed to the mobile computing device, causing one or more motors to complete moving the first enclosure away from the second housing enclosure.
[0075] In some implementations, the method may further include, when the one or more motors have completed moving the mobile computing device toward the user's location, causing, based on the verbal utterance and / or another verbal utterance, one or more third motors of the mobile computing device to move the second housing enclosure away from the third housing enclosure of the mobile computing device and to move a camera of the mobile computing device toward the user. In some implementations, the method may further include, when the one or more motors have completed moving the mobile computing device toward the user's location, causing, based on the verbal utterance and / or another verbal utterance, one or more fourth motors to rotate the first housing enclosure about an axis that intersects a surface of the third housing enclosure of the mobile computing device to facilitate orienting the display panel toward the user. [Explanation of symbols]
[0076] 102, 204, 304, 404 Mobile computing devices 104 Display Panel 106 first housing enclosure 108 Second housing enclosure 110 3rd housing enclosure 112 Camera 114, 116 Microphones 124, 136 Angle of separation 126 Arm, Another Arm 128 Rotatable Plate 132 Axis 134 Another Axis 202, 302, 402 Users Rooms 206, 306, and 406 208, 308, 408 Couch 210, 310 Oral utterances 212 Response output 312 Routes 314 Television 410 Oral utterances, subsequent oral utterances 412 First Part 414 Second Part 500 Systems 502 Server Device 504 Automated Assistant 506 Input Processing Engine 508 Audio Processing Module 510 Data Parsing Module 512 Parameter Module 514 Output Generation Engine 516 Assistant Data 518 Computing Devices 520 Assistant Interface 522 Client Automated Assistant 524 Spatial Processing Engine 526 Power supply 528 Electric Engine 530 Location Engine 532 Motor Control Engine 534 Content Rendering Engine 536 Network 538 Client Data 710 Computer Systems 712 Bus Subsystem 714 Processor 716 Network Interface Subsystem 720 User Interface Output Devices 722 User Interface Input Devices 724 Memory Subsystem 725 Memory 726 File Storage Subsystem 730 Main Random Access Memory (RAM) 732 Read-Only Memory (ROM)
Claims
1. 1. A method implemented by one or more processors of a computing device, comprising: based on processing audio data detected via a plurality of microphones of the computing device; A user provides a verbal utterance directed to an automated assistant operating at least in part on the computing device; and and a location of the user providing the spoken utterance relative to the computing device, determining that the plurality of microphones are mounted in a given housing enclosure of the computing device; In response to determining that the user has provided the oral utterance, activating a motor included in the given housing enclosure that includes the plurality of microphones to orient a display panel of the computing device toward the determined position of the user; the display panel is within a display panel housing enclosure; the display panel housing enclosure is separate from but coupled to the given housing enclosure; activating the motor included in the given housing enclosure to rotate the display panel housing enclosure via the coupling to the given housing enclosure about an axis of the motor; causing the display panel of the computing device to render graphical content generated by the automated assistant and responsive to the verbal utterance; the graphical content is rendered by the display panel at least after the orientation is directed to the determined position of the user; identifying a distance condition dependent on whether a response to the verbal utterance includes any graphical content; determining that the location of the user does not satisfy the distance condition for the computing device; activating one or more wheel motors each driving one or more corresponding wheels coupled to the given housing enclosure to cause the computing device to navigate closer to the user; actuating the one or more wheel motors to cause the computing device to navigate closer to the user, further in response to determining that the position does not satisfy the distance condition; activating, following the computing device navigating closer to the user, the orientation of the display panel is directed towards the user to facilitate rendering the graphical content towards the user; Including, method.
2. determining that the response to the verbal utterance includes the graphical content, and determining that actuating the motor to orient the display panel toward the determined position of the user is responsive to determining that the response of the verbal utterance includes the graphical content. The method of claim 1.
3. 2. The method of claim 1, wherein the coupling of the display panel housing enclosure to the given housing enclosure is via a fulcrum, and the display panel housing enclosure is further adjustable about a fulcrum axis of the fulcrum.
4. The method of claim 1 , wherein the distance condition is determined based on the graphical content.
5. determining that additional verbal utterances are being provided while actuating the motor to orient the display panel toward the determined position of the user; and transitioning the motor to a reduced power state in response to determining that the additional verbal utterance has been provided while the motor is operating. The method of claim 1.
6. A given housing enclosure; a plurality of microphones mounted in the given housing enclosure; a wheel coupled to the given housing enclosure; wheel motors within the given housing enclosure, each of the wheel motors driving a respective one of the wheels; a display panel housing enclosure separate from but coupled to the given housing enclosure; and a motor having a motor shaft, the motor within the given housing enclosure and operable to rotate the display panel housing enclosure relative to the given housing enclosure about the motor shaft; A memory for storing instructions; and one or more processors for executing the instructions, the instructions configuring the one or more processors to: based on processing the audio data detected via the microphone; A user provides a verbal utterance directed to an automated assistant operating at least in part on the computing device; and a location of the user providing the spoken utterance relative to the computing device; and and determining In response to determining that the user has provided the oral utterance, activating the motor to orient a display panel toward the determined position of the user; and causing the display panel to render graphical content generated by the automated assistant and responsive to the verbal utterance; wherein the graphical content is rendered by the display panel at least after the orientation is directed to the determined position of the user; and determining that the location of the user does not satisfy a distance condition relative to the computing device and that depends on whether the spoken utterance response includes any graphical content; activating the wheel motors to cause the computing device to navigate closer to the user; In operating the wheel motors to cause the computing device to navigate closer to the user, the one or more processors, in response to determining that the position does not satisfy the distance condition, further operate the wheel motors; and actuating, subsequent to the computing device navigating closer to the user, the orientation of the display panel is directed towards the user to facilitate rendering the graphical content towards the user; To carry out Computing device.
7. The one or more processors further, in executing the instructions, determining that the response to the verbal utterance includes the graphical content; and wherein in actuating the motor, the one or more processors further actuate the motor in response to determining that the response to the verbal utterance includes the graphical content, to orient the display panel toward the determined position of the user. The computing device of claim 6.
8. 7. The computing device of claim 6, further comprising a fulcrum coupling the display panel housing enclosure to the given housing enclosure, the display panel housing enclosure being further adjustable about a fulcrum axis of the fulcrum.
9. The computing device of claim 6, wherein one or more of the processors determine the distance condition based on the graphical content.
10. The one or more processors further, in executing the instructions, before the user provides the oral utterance and while the user is in the same position relative to the computing device; based on processing previous audio data detected via the microphone of the computing device; A user has provided a previous verbal utterance directed to the automated assistant; and determining that the audio-only content is responsive to the prior verbal utterance; in response to determining that the user provided the previous verbal utterance and based on that the audio-only content responded to the previous verbal utterance; causing the audio-only content to be rendered through at least one speaker of the computing device and independent of any actuation of the motor. To do that, The computing device of claim 6.
Citation Information
Patent Citations
Program, control method, and information communication device
JP2018137744A
Robots for interactive comedy and companionship
WO2018045081A1
Social robot with environmental control feature
WO2018089700A1
Transmission method, transmission device, and program
WO2018110373A1