An electrically powered computing device that automatically adjusts the device's position and / or orientation of an interface in response to an automated assistant request
By using a motor on a mobile computing device to adjust the angle and navigation of the display panel, the problem of users having difficulty accessing the computing device when they are in dangerous locations is solved, achieving secure content rendering and power savings.
Patent Information
- Application Number
- CN201980091797.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-04-29
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2039-04-29
AI Technical Summary
When a user performs certain tasks, it is difficult to access the computing device reasonably or securely to obtain information, especially when the user is in a dangerous position, it is impossible to effectively perceive or hear what the computing device provides.
The mobile computing device adjusts the angle of the display panel through a motor, navigates selectively according to the user's relative position to render the content, and bypasses the navigation when the user requests audio content to save power.
It realizes that when users cannot access the computing device safely, the display panel is automatically adjusted so that users can view content, and save power and avoid delays when navigation is not required.
Smart Images

Figure CN113424124B_ABST
Abstract
Description
Background Art
[0001] A human may engage in a human-computer dialogue with an interactive software application, referred to herein as an "automated assistant" (also referred to as a "digital agent," "chat program," "interactive personal assistant," "intelligent personal assistant," "conversational agent," etc.). For example, a human (who may be referred to as a "user" when they interact with an automated assistant) may provide commands and / or requests using spoken natural language input (i.e., utterances) and / or by providing textual (e.g., typed) natural language input, where the spoken natural language input may in some cases be converted to text and then processed. Although the use of an automated assistant can allow for easier access to information and a more convenient means for controlling peripheral devices, perceiving display content and / or audio content can be difficult in some cases.
[0002] For example, when a user is engrossed in certain tasks in a room in their home and expects to obtain helpful information about the tasks via a computing device in a different room, the user may not be able to reasonably and / or safely access the computing device. This can be particularly evident when the user is performing skilled labor and / or otherwise consuming energy to be in their current situation (e.g., standing on a ladder, working under a vehicle, painting their home, etc.). If a user requests certain information while working in this situation, the user may not be able to properly hear and / or see the rendered content. For example, although a user may have a computing device in their garage, in order to view helpful content while working in their garage, the user may not be able to see the display panel of the computing device when standing in certain locations within the garage. As a result, the user may unfortunately need to pause the progress of their work in order to view the display panel for perceiving the information provided via the computing device. In addition, depending on the amount of time to render the content, a particular user may not have time to properly perceive the content—assuming that they must first navigate around various fixtures and / or people to reach the computing device. In this case, if the user does not have a chance to perceive the content, the user may end up having to re-request the content, thereby wasting computing resources and power of the computing device. Summary of the invention
[0003] The embodiment described herein relates to a mobile computing device, which selectively navigates to a user for rendering certain content to the user, and adjusts the viewing angle of the display panel of the mobile computing device according to the relative position of the user. The mobile computing device can include multiple parts (i.e., housing packaging), and each part can include one or more motors for adjusting the physical position and / or arrangement of one or more parts. As an example, a user can provide a verbal utterance, such as "Assistant, what is my schedule for today?" ("Assistant, what is my schedule for today"), and in response, the mobile computing device can navigate toward the position of the user, adjust the angle of the display panel of the mobile computing device, and render the display content of the arrangement characterizing the user. However, when the user requests audio content (e.g., audio content without corresponding display content), the mobile computing device can bypass navigation to the user, and the mobile computing device determines that the current position of the mobile computing device corresponds to the audio content and will be audible to the user. In the case where the mobile computing device is not within the distance for rendering audible audio content, the mobile computing device can determine the position of the user and navigate toward the user, at least until the mobile computing device is within a specific distance for rendering audible audio content.
[0004] In order for the mobile computing device to determine how to operate the motor of the mobile computing device to arrange the part of the mobile computing device for rendering content, the mobile computing device can process data based on one or more sensors. For example, the mobile computing device can include one or more microphones (e.g., microphone arrays) that respond to sounds originating from different directions. One or more processors can process the output from the microphone to determine the origin of the sound relative to the mobile computing device. In this way, if the mobile computing device determines that the user has requested content that should be rendered closer to the user, the mobile computing device can navigate to the user's location, as determined based on the output from the microphone. Allowing the mobile computing device to operate in this way can provide relief for impaired users who may not be able to efficiently navigate to computing devices for information and / or other media. In addition, because the mobile computing device can make a determination about when to navigate to the user and when not to navigate to the user, at least when to render content, the mobile computing device can save power and other computing resources. For example, if the mobile computing device navigates to the user indiscriminately with respect to the type of content to be rendered, the mobile computing device may consume more power to navigate to the user than to render content exclusively without navigating to the user. Furthermore, rendering content without first navigating the user when there is no need to do so can avoid unnecessary delays in the rendering of the content.
[0005] In some embodiments, a mobile computing device can include a top housing package that includes a display panel that can have a viewing angle that is adjustable via a top housing package motor. For example, when the mobile computing device determines that a user has provided a spoken utterance and the viewing angle of the display panel should be adjusted so that the display panel is pointed toward the user, the top housing package motor can adjust the position of the display panel. When the mobile computing device has completed rendering of particular content, the top housing package motor can manipulate the display panel back to a resting position, which can consume less space than when the display panel has been adjusted toward the direction of the user.
[0006] In some embodiments, the mobile computing device can also include an intermediate housing package and / or a bottom housing package, which can each accommodate one or more parts of the mobile computing device. For example, the intermediate housing package and / or the bottom housing package can include one or more cameras for capturing images of the surrounding environment of the mobile computing device. Image data generated based on the output of the camera of the mobile computing device can be used to determine the location of the user so as to allow the mobile computing device to navigate toward the location when the user provides certain commands to the mobile computing device. In some embodiments, when the mobile computing device is operating in a sleep mode, the intermediate housing package and / or the bottom housing package can be positioned below the top housing package; and the intermediate housing package and / or the bottom housing package can include one or more motors for rearranging the mobile computing device including the top housing package when switching out of the sleep mode. For example, in response to the user providing an invocation phrase such as "Assistant..." ("Assistant..."), the mobile computing device switches out of the sleep mode and enters the operating mode. During the conversion, the one or more motors of the mobile computing device can cause the camera and / or the display panel to be directed toward the user. In some embodiments, a first group of motors in the one or more motors can control the orientation of the camera, and a second group of motors in the one or more motors can control the individual orientation of the display panel. In this way, the cameras can have a separate orientation relative to the orientation of the display panel. Additionally or alternatively, the cameras can have the same orientation relative to the orientation of the display panel based on the movement of one or more motors.
[0007] Furthermore, during the transition from the compact arrangement of the mobile computing device to the expanded arrangement of the mobile computing device, the one or more microphones of the mobile computing device can monitor for additional input from the user. When additional input is provided by the user during operation of the one or more motors of the mobile computing device, noise from the motors can interrupt certain frequencies of the verbal input from the user. Therefore, in order to eliminate the negative impact on the quality of the sound captured by the microphones of the mobile computing device, the mobile computing device can modify and / or pause the operation of the one or more motors of the mobile computing device while the user provides subsequent verbal utterances.
[0008] For example, after the user provides the invocation phrase "assistant," and while the motors of the mobile computing device operate to extend the display panel in the direction of the user, the user can provide a subsequent spoken utterance. The subsequent spoken utterance can be "…show the security camera in front of the house," which can cause the automated assistant to invoke an application for viewing a live streaming video from a security camera. However, because the motors of the mobile computing device are operating while the user provides the subsequent spoken utterance, the mobile computing device can determine that the user is providing the spoken utterance, and in response, modify one or more operations of one or more motors of the mobile computing device. For example, the mobile computing device can stop the operation of the motor that extends the display panel in the direction of the user so as to eliminate motor noise that would interrupt the mobile computing device when generating audio data representing the spoken utterance. When the mobile computing device determines that the spoken utterance is no longer provided by the user and / or is otherwise completed, the operation of the one or more motors of the mobile computing device can continue.
[0009] Each motor of the mobile computing device can perform various tasks to achieve certain operations of the mobile computing device. In some embodiments, the bottom housing package of the mobile computing device can include one or more motors connected to one or more wheels (e.g., cylindrical wheel(s), spherical wheel(s), Mecanum wheel(s), etc.) for navigating the mobile computing device to a location. The bottom housing package can also include one or more other motors for manipulating the middle housing package of the mobile computing device (e.g., rotating the middle housing package around an axis perpendicular to the surface of the bottom housing package).
[0010] In some implementations, the mobile computing device can perform one or more different responsive gestures to indicate that the mobile computing device is receiving input from the user, thereby confirming the input. For example, in response to detecting a spoken utterance from the user, the mobile computing device can determine whether the user is within the viewing range of a camera of the mobile computing device. If the user is within the viewing range of the camera, the mobile computing device can operate one or more motors to invoke physical motion by the mobile computing device, thereby indicating to the user that the mobile computing device is confirming the input.
[0011] The above description is provided as an overview of some embodiments of the present disclosure. Further descriptions of these embodiments and other embodiments are described in more detail below.
[0012] Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform methods such as one or more of the methods described above and / or elsewhere herein. Other embodiments may include a system of one or more computers and / or one or more robots including one or more processors operable to execute the stored instructions to perform methods such as one or more of the methods described above and / or elsewhere herein.
[0013] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1A , Figure 1B and Figure 1C Illustrated is a view of a mobile computing device that automatically and selectively navigates in response to commands from one or more users.
[0015] Figure 2 A view illustrating a user providing a spoken utterance to a mobile computing device in order to invoke a response from an automated assistant accessible via the mobile computing device.
[0016] Figure 3 The illustration shows a user providing spoken words that navigate the mobile computing device toward the user and arrange different case enclosures of the mobile computing device to protect graphical content toward the user.
[0017] Figure 4 Illustrated is a view of a user providing spoken utterances to a mobile computing device, which can pause intermittently during navigation to capture any additional spoken utterances from the user.
[0018] Figure 5 Further illustrated is a system for operating a computing device that provides access to an automated assistant and automatically moves toward and / or away from a user in response to spoken utterances.
[0019] Figure 6 Illustrated is a method for rendering content at a mobile computing device that selectively and automatically navigates to a user in response to spoken utterances. DETAILED DESCRIPTION
[0020] Figure 1A , 1B1C illustrates a view of a mobile computing device 102 that automatically and selectively navigates in response to commands from one or more users. Specifically, Figure 1A A perspective view 100 of a mobile computing device 102 in a folded state is illustrated, wherein the housing enclosures of the mobile computing device 102 can be closest to each other. In some implementations, the mobile computing device 102 can include one or more housing enclosures, including but not limited to one or more of a first housing enclosure 106, a second housing enclosure 108, and / or a third housing enclosure 110. One or more housing enclosures can include one or more motors for maneuvering a particular housing enclosure to a particular position and / or toward a particular destination.
[0021] In some embodiments, the first housing package 106 can include one or more first motors (i.e., a single motor or multiple motors) that operate to adjust the orientation of the display panel 104 of the mobile computing device 102. The first motor of the first housing package 106 can be controlled by one or more processors of the mobile computing device 102 and can be powered by a portable power source of the mobile computing device 102. The portable power source can be a rechargeable power source, such as a battery and / or a capacitor, and the power management circuit of the mobile computing device 102 can adjust the output of the portable power source for providing power to the first motor and / or any other motors of the mobile computing device 102. During operation of the mobile computing device 102, the mobile computing device 102 can receive input from a user and determine a response to provide to the user. When the response includes display content, the first motor can adjust the display panel 104 to point toward the user. For example, as Figure 1B As shown in view 120 of FIG. 1 , the display panel 104 can be manipulated in a direction that increases the separation angle 124 between the first housing enclosure 106 and the second housing enclosure 108. However, when the response includes audio content, the mobile computing device 102 can remain in the compressed state (e.g., when the response includes audio content) without providing corresponding display content. Figure 1A ), while the first motor does not adjust the orientation of the display panel 104.
[0022] In some embodiments, the second housing enclosure 108 can include one or more second motors for manipulating the orientation of the first housing enclosure 106 and / or the second housing enclosure 108 relative to the third housing enclosure 110. For example, Figure 1C As shown in view 130 in FIG. 1 , the one or more second motors can include a motor that manipulates the first housing package 106 around an axis 132 that is perpendicular to the surface of the second housing package 108. Thus, the first motor included in the first housing package 106 can modify the separation angle between the first housing package 106 and the second housing package 108, and the second motor included in the second housing package 108 can rotate the first housing package 106 around the axis 132.
[0023] In some embodiments, the mobile computing device 102 can include a third housing package 110 that includes one or more third motors to manipulate the first housing package 106, the second housing package 108, and / or the third housing package 110, and / or also navigate the mobile computing device 102 to one or more different destinations. For example, the third housing package 110 can include one or more motors for manipulating the second housing package 108 around another axis 134. The other axis 134 can intersect with a rotatable plate 128 that can be attached to the third motor encapsulated within the third housing package 110. In addition, the rotatable plate 128 can be connected to an arm 126 on which the second housing package 108 can be mounted. In some embodiments, a second motor encapsulated within the second housing package 108 can be connected to the arm 126, which can operate as a fulcrum to allow adjustment of a separation angle 136 between the third housing package 110 and the second housing package 108.
[0024] In some embodiments, the mobile computing device 102 can include another arm 126 that can operate as another fulcrum or other means to assist the second motor in adjusting the separation angle 124 between the first housing package 106 and the second housing package 108. The arrangement of the mobile computing device 102 can depend on the situation in which the mobile computing device 102 and / or another computing device receives input. For example, in some embodiments, the mobile computing device 102 can include one or more microphones oriented in different directions. For example, the mobile computing device 102 can include a first set of one or more microphones 114 oriented in a first direction, and a second set of one or more microphones 116 oriented in a second direction different from the first direction. The microphone can be attached to the second housing package 108 and / or any other housing package or combination of housing packages.
[0025] When the mobile computing device 102 receives input from the user, the signals from the one or more microphones can be processed to determine the position of the user relative to the mobile computing device 102. When the position is determined, the one or more processors of the mobile computing device 102 can cause one or more motors of the mobile computing device 102 to align the housing package so that the second set of microphones 116 is directed toward the user. In addition, in some embodiments, the mobile computing device 102 can include one or more cameras 112. The camera 112 can be connected to the second housing package and / or any other housing package or combination of housing packages. In response to the mobile computing device 102 receiving the input and determining the position of the user, the one or more processors can cause one or more motors to align the mobile computing device 102 so that the camera 112 is directed toward the user.
[0026] As a non-limiting example, the mobile computing device 102 can determine that the child has indicated a verbal utterance at the mobile computing device 102 while the mobile computing device 102 is on top of a table that is taller than the child. In response to receiving the verbal utterance, the mobile computing device 102 can use signals from one or more of the microphone 114, the microphone 116, and / or the camera 112 to determine the position of the child relative to the mobile computing device 102. In some implementations, the processing of the audio and / or video signals can be offloaded to a remote device such as a remote server via a network to which the mobile computing device 102 is connected. The mobile computing device 102 can determine based on the processing that an anatomical feature (e.g., an eye, an ear, a face, a mouth, and / or any other anatomical feature) is located below the table. Accordingly, in order to point the display panel 104 toward the user, one or more motors of the mobile computing device 102 can align the display panel 104 in a direction below the table.
[0027] Figure 2 View 200 illustrates user 202 providing spoken utterances to mobile computing device 204 in order to invoke a response from an automated assistant accessible via mobile computing device 204. Specifically, user 202 can provide spoken utterances 210, which can be captured via one or more microphones of mobile computing device 204. Mobile computing device 204 and / or a computing device in communication with mobile computing device 204 can process audio data representing spoken utterances 210. Based on the processing, mobile computing device 204 can determine one or more actions that user 202 is requesting. Further, mobile computing device 204 can determine whether performance of the one or more actions involves rendering graphical content at a display panel of mobile computing device 204. When the one or more actions do not involve rendering graphical content for user 202, in some implementations, mobile computing device 204 can navigate around obstacles, such as recliner 208, within room 206 in which user 202 is located.
[0028] Instead, mobile computing device 204 can determine that one or more actions involve rendering audio content, and further determine whether user 202 is located within a distance for effectively rendering audible content to user 202. In other words, mobile computing device 204 can determine whether user 202 is located close enough to mobile computing device 204 to hear any audio output generated at mobile computing device 204. If user 202 is within a distance to audibly render the audio content, mobile computing device 204 can navigate around closer to user 202. However, if the user is not within a distance to audibly render the audio content, mobile computing device 204 can control one or more motors of mobile computing device 204 for navigating closer to user 202.
[0029] When the mobile computing device 204 reaches a distance to audibly render the audio content for the user 202 or otherwise determines that the mobile computing device 204 is already within a distance to audibly render the audio content, the mobile computing device 204 can provide a response output 212. For example, when the spoken utterance includes natural language content, such as "Assistant, what's my schedule for today?", the automated assistant can provide a response output, such as "Okay, you are meeting Lenny for coffee at 9:00 AM, and you are playing a game with Sol and Darren at 11:30 AM." In this way, power can be saved at the mobile computing device 204 by selectively navigating or not navigating to the user 202, depending on one or more actions requested by the user 202. Furthermore, the ability of mobile computing device 204 to navigate to user 202 and provide responses when user 202 is unable to reach mobile computing device 204 can eliminate the need for user 202 to stop what they are doing in certain situations.
[0030] Figure 3 The view 300 illustrates a user 302 providing a spoken utterance 310 that causes a mobile computing device 304 to navigate toward the user 302 and arrange different housing packages of the mobile computing device 304 to protect graphical content toward the user 302. The mobile computing device 304 can selectively navigate toward the user 302 depending on whether the user 302 is requesting that a particular type of content be provided to the user 302. For example, Figure 2 , when the user has requested audio content and the mobile computing device 204 has determined that the user 202 is within a threshold distance for audibly rendering the audio content for the user 202, the mobile computing device 204 can bypass navigating to the user 202. However, when the user requests graphical content, such as in Figure 3 , mobile computing device 304 is able to navigate user 302 to present graphical content at a display panel of mobile computing device 304 .
[0031] For example, Figure 3As provided in FIG. 3 , user 302 can provide spoken utterance 310, such as “Assistant, show me security video from yesterday”. In response to receiving spoken utterance 310, mobile computing device 304 can generate audio data representing the spoken utterance, and process the audio data at mobile computing device 304 and / or send the audio data to another computing device for processing. Based on the processing of the audio data, mobile computing device 304 can determine that user 302 is requesting mobile computing device 304 and / or any other device with display functionality to provide playback of security video for user 302. Based on this determination, mobile computing device 304 can determine the location of user 302 relative to mobile computing device 304. Additionally or alternatively, mobile computing device 304 can identify one or more obstacles present in room 306 that mobile computing device 304 navigates to in order to reach user 302. For example, using image data captured via a camera of mobile computing device 304, it can be determined that recliner 308 is separating user 302 from mobile computing device 304. Using the image data, mobile computing device 304 and / or a remote computing device processing the image data can generate route 312 for reaching user 302 .
[0032] In some implementations, mobile computing device 304 can be connected to a local network to which other computing devices are connected. When mobile computing device 304 determines that mobile computing device 304 is unable to navigate to user 302 due to one or more obstacles and / or is unable to navigate to the user within a threshold amount of time, an automated assistant accessible via mobile computing device 304 can identify other display-enabled devices connected via the local area network. For example, the automated assistant can determine that television 314 is located in the same room as user 302, and also determine that mobile computing device 304 is unable to reach the user. Based on these determinations, the automated assistant can cause the requested display content to be rendered at television 314 and / or any other computing device located in room 306 and enabled for display.
[0033] However, when the mobile computing device 304 is able to navigate the route 312 to the user 302, the mobile computing device 304 may identify one or more anatomical features of the user 302 when it reaches the user 302 and / or when the mobile computing device 304 is en route to the location of the user 302. For example, when the mobile computing device 304 is navigating the route 312 to the user 302, the mobile computing device 304 is able to determine that the user 302 is within the viewing window of the camera of the mobile computing device 304. In response to this determination, the mobile computing device 304 is able to use the image data captured via the camera to identify the eyes, mouth, and / or ears of the user 302. Based on identifying one or more of these anatomical features of the user 302, the mobile computing device 304 is able to cause one or more motors of the mobile computing device 304 to manipulate the display panel of the mobile computing device 304 in the direction of the user 302.
[0034] For example, when the display panel is connected to the first housing package of the mobile computing device 304, the one or more first motors of the mobile computing device 304 can increase the separation angle between the first housing package and the second housing package. In addition, based on the determined position of the anatomical features of the user 302 relative to the mobile computing device 304, the one or more motors of the mobile computing device 304 can further increase the height of the mobile computing device 304 so that the display panel is easier to view by the user 302. For example, the one or more second motors of the mobile computing device 304 can cause the second housing package of the mobile computing device 304 to have an increased separation angle relative to the third housing package of the mobile computing device 304. These increases in separation angles can cause the mobile computing device 304 to transition from being in a compressed state to being in an expanded state, thereby increasing the height of the mobile computing device 304. When the mobile computing device 304 has finished rendering the graphical content according to the spoken utterance 310, the mobile computing device 304 can return to the folded state in order to preserve the stored energy of the rechargeable power supply of the mobile computing device 304.
[0035] Figure 4A view 400 of a user 402 providing spoken utterances to a mobile computing device 404 is illustrated, the mobile computing device being able to intermittently pause during navigation to capture any additional spoken utterances from the user 402. The mobile computing device 404 is able to and provides access to an automated assistant that is able to respond to a variety of different inputs from one or more users. The user 402 is able to provide spoken utterances that are able to be processed at the mobile computing device 404 and / or another computing device associated with the mobile computing device 404. The mobile computing device 404 is able to include one or more microphones that are able to provide output signals in response to spoken inputs from the user. In order to eliminate noise that may otherwise affect the spoken input, the mobile computing device 404 is able to determine whether the user 402 is providing spoken input while one or more motors of the mobile computing device 404 are operating. In response to determining that spoken utterances are being provided while one or more motors are operating, the mobile computing device 404 is able to cause the one or more motors to enter a low power state in order to reduce the amount of noise generated by the one or more motors.
[0036] For example, user 402 can provide spoken utterances 410, such as "assistant, video call my brother...", and mobile computing device 404 can receive spoken utterances 410 and determine that spoken utterances 410 correspond to actions involving a camera and / or rendering graphical content. In response to determining that the actions involve a camera of mobile computing device 404 and rendering graphical content, mobile computing device 404 can navigate toward the location of user 402. As mobile computing device 404 traverses a first portion 412 of a route, user 402 can provide subsequent spoken utterances 410, such as "... And also, secure the alarm for the house." While traversing the first portion 412 of the route, mobile computing device 404 can determine that user 402 is providing subsequent spoken utterances. In response, mobile computing device 404 can cause one or more motors of mobile computing device 404 to enter a lower power state relative to the power state in which one or more motors were operating when mobile computing device 404 was traversing the first portion 412 of the route. For example, one or more motors can pause their respective operations, thereby causing the mobile computing device 404 to pause after the first portion 412 of the route of the user 402 .
[0037] The mobile computing device 404 and / or the automated assistant determines that the subsequent spoken utterance has been completed and / or is otherwise no longer directed to the mobile computing device 404, and the mobile computing device 404 can proceed to the second portion 414 of the traversal route toward the location of the user 402. The second portion 414 of the route can include navigating through a room 406 that includes a recliner 408 and / or other obstacles. The mobile computing device 404 can use a camera to identify such obstacles and the user 402 with the user's prior permission. In some embodiments, while the mobile computing device 404 is traversing the second portion of the route 414, the automated assistant can initiate the performance of other actions requested by the user 402. Specifically, while the mobile computing device 404 is traversing the second portion 414 of the route toward the location of the user 402, the mobile computing device 404 can initialize the alarm for the house. The action can be initialized based on the determination and the action does not involve rendering graphical content that the user wants to see and / or does not involve rendering audio data that the mobile computing device 404 attempts to render at a distance that allows the audio content to be audible to the user 402.
[0038] Figure 5 Illustrated is a system 500 for operating a computing device 518 to selectively navigate to a user for rendering certain content to the user and switching motor operation depending on whether the user is providing a spoken utterance to the computing device 518. Automated assistant 504 can operate as part of an assistant application provided at one or more computing devices, such as computing device 518 and / or server device 502. The user can interact with automated assistant 504 via an assistant interface, which can be a microphone, camera, touch screen display, user interface, and / or any other device capable of providing an interface between a user and an application.
[0039] For example, a user can initialize the automated assistant 504 by providing verbal, textual, and / or graphical input to the assistant interface to cause the automated assistant 504 to perform a function (e.g., provide data, control peripherals, access agents, generate input and / or output, etc.). The computing device 518 can include a display device, which can be a display panel including a touch interface for receiving touch input and / or gestures to allow the user to control the application of the computing device 518 via the touch interface. In some embodiments, the computing device 518 can lack a display device, thereby providing an audible user interface output without providing a graphical user interface output. In addition, the computing device 518 can provide a user interface such as a microphone for receiving spoken natural language input from the user. In some embodiments, the computing device 518 can include a touch interface and can be without a camera, but can optionally include one or more other sensors.
[0040] Computing device 518 and / or other computing devices can communicate with server device 502 via network 536, such as the Internet. Additionally, computing device 518 and other computing devices can communicate with each other via a local area network (LAN), such as a WiFi network. Computing device 518 can offload computing tasks to server device 502 to conserve computing resources at computing device 518. For example, server device 502 can host automated assistant 504, and computing device 518 can send input received at one or more assistant interfaces 420 to server device 502. However, in some implementations, automated assistant 504 can be hosted at computing device 518 as a client automated assistant 522.
[0041] In various embodiments, all or less than all aspects of automated assistant 504 can be implemented on computing device 518. In some of those embodiments, aspects of automated assistant 504 are implemented via client automated assistant 522 of computing device 518 and interface with server device 502 that implements other aspects of automated assistant 504. Server device 502 can optionally serve multiple users and their associated assistant applications via multiple threads. In embodiments where all or less than all aspects of automated assistant 504 are implemented via client automated assistant 522 at computing device 518, client automated assistant 522 can be an application separate from (e.g., installed “on top of”) the operating system of computing device 518—or alternatively implemented directly by (e.g., treated as an application of, but integrated with) the operating system of computing device 518.
[0042] In some implementations, automated assistant 504 and / or client automated assistant 522 can include an input processing engine 506 that can employ a number of different modules to process input and / or output of computing device 518 and / or server device 502. For example, input processing engine 506 can include a speech processing module 508 that can process audio data received at assistant interface 420 to identify text contained in the audio data. The audio data can be sent from, for example, computing device 518 to server device 502 in order to conserve computing resources at computing device 518.
[0043] The process for converting audio data into text can include a speech recognition algorithm, which can use a neural network and / or a statistical model to identify an audio data group corresponding to a word or phrase. The text converted from the audio data can be parsed by a data parsing module 510, and is available to an automated assistant as a text data that can be used to generate and / or identify a command phrase from a user. In some embodiments, the output data provided by the data parsing module 510 can be provided to a parameter module 512 to determine whether the user has provided an input corresponding to a specific action and / or routine that can be performed by an automated assistant 504 and / or an application or agent that can be accessed by an automated assistant 504. For example, assistant data 516 can be stored in a server device 502 and / or a computing device 518 as client data 538, and can include data defining one or more actions that can be performed by an automated assistant 504 and / or a client automated assistant 522, and parameters necessary for performing these actions.
[0044] Figure 5 Further illustrated is a system 500 for operating a computing device 518 that provides access to an automated assistant 504 and automatically moves toward and / or away from a user in response to a spoken utterance. The computing device 518 can be powered by one or more power supplies 526 that can be rechargeable and / or can allow the computing device 518 to be portable. A motor control engine 532 can be powered by the power supply 526 and determine when to control one or more motors of the computing device 518. For example, the motor control engine 532 can determine one or more operating states of the computing device 518 and control one or more motors of the computing device 518 to reflect the one or more operating states. For example, when the computing device 518 has received a spoken utterance at the assistant interface 520 of the computing device 518, the motor control engine 532 can determine that the user is providing a spoken utterance and cause one or more motors to operate to facilitate indicating that the computing device 518 is confirming the spoken utterance. The one or more motors can, for example, cause the computing device 518 to shake and / or dance when the spoken utterance is received. Alternatively or additionally, the motor control engine 532 can cause one or more motors to maneuver the computing device 518 back and forth via the ground wheels of the computing device 518 to indicate that the computing device 518 is downloading and / or uploading data over a network. Alternatively or additionally, the motor control engine 532 can cause one or more motors to arrange the housing package of the computing device 518 to be in a compressed or relaxed state, which indicates that the computing device 518 is operating in a low power mode and / or a sleep mode.
[0045] When operating in sleep mode, the computing device 518 can monitor the call phrases spoken by the user, and / or can perform voice activity detection. When the computing device is performing voice activity detection, the computing device 518 can determine whether the input to the microphone corresponds to a human. In addition, when the computing device 518 is controlled by one or more motors of the computing device 518, voice activity detection can be performed. In some embodiments, the threshold for determining whether a human voice has been detected can include a threshold when one or more motors are operating, and another threshold when one or more motors are not operating. For example, when the computing device 518 is in sleep mode, voice activity detection can be performed according to a first threshold, which can satisfy us when the first percentage of the incoming noise corresponds to human voice. However, when the computing device 518 is in wake-up mode, voice activity detection can be performed according to a second threshold, which can be satisfied when the second percentage of the incoming noise corresponds to human voice, and the second percentage of the incoming noise is higher than the first percentage of the incoming noise. Additionally, when computing device 518 is in wake mode and one or more motors are operating to realign computing device 518 and / or navigate computing device 518, voice activity detection can be performed according to a third price hold, which can be satisfied when a third percentage of incoming noise corresponds to human speech. The third percentage of incoming route can be greater than and / or equal to the second percentage of incoming noise and / or the first percentage of incoming noise.
[0046] In some implementations, in response to the computing device 518 determining that human speech has been detected, the spatial processing engine 524 processes incoming data from one or more sensors to determine where the human speech is coming from. Spatial data characterizing the location of a source of human speech, such as a user, can be generated by the spatial processing engine 524 and transmitted to the location engine 530. The location engine 530 can use the spatial data to generate a route for navigating the computing device 518 from my current location at the computing device 518 to the location of the source of human speech. Route data can be generated by the location engine 530 and transmitted to the motor control engine 532. The motor control engine 532 can use the route data to control one or more motors of the computing device 518 for navigating to the location of the user and / or the source of human speech.
[0047] In some embodiments, the spatial processing engine 524 can process incoming data from one or more sensors of the computing device 518 to determine whether the computing device 518 is located within a distance from a human voice source for rendering audible audio. For example, an automated assistant can receive a request from a user and can determine one or more actions that the user is requesting. One or more actions can be transmitted to the content rendering engine 534, which can determine whether the user is requesting audio content to be rendered, graphics content to be rendered, and / or audio or graphics content to be rendered. In response to determining that the user has requested audio content to be rendered, the spatial processing engine 524 can determine whether the computing device 518 is located within a distance from the user for generating audio content audible to the user. When the computing device 518 determines that the computing device 518 is not within a distance for generating audible content, the computing device 518 can control one or more motors to navigate the computing device 518 to within a distance for generating audible content.
[0048] Alternatively, in response to determining that the user has requested that the graphical content be rendered, the spatial processing engine 524 can determine whether the computing device 518 is located within another distance from the user for generating graphical content that will be visible to the user. When the computing device 518 determines that the computing device 518 is not within another distance for generating visible graphical content, the computing device 518 can control one or more motors to navigate the computing device 518 to another distance for generating visible graphical content. In some embodiments, when the computing device 518 is rendering graphical content, the amount of distance between the user and the computing device 518 can be based on specific attributes of the graphical content. For example, when the graphical content includes text of size X, the computing device 518 navigates to within a distance of N from the user. However, when the graphical content includes text of size Y less than size X, the computing device 518 can navigate to within a distance of N from the user, where and less than M. Alternatively or in addition, the distance between the user and the computing device 518 can be based on the type of content to be generated. For example, computing device 518 can navigate to a distance of H from the user when the graphics content to be rendered includes video content, and can navigate to a distance of K from the user when the graphics content to be rendered includes a static image, where H is less than K.
[0049] Figure 6The method 600 for rendering content at a mobile computing device is illustrated, and the mobile computing device selectively and automatically navigates to a user in response to a spoken utterance. The method 600 can be performed by one or more computing devices, applications, and / or any other device or module that can respond to a spoken utterance. The method 600 can include an operation 602 of determining whether a spoken utterance from a user has been received. The mobile computing device can include one or more microphones, and the mobile computing device can use the one or more microphones to detect a verbal input from a user. In addition, the mobile computing device can provide access to an automated assistant, which can initialize an action and / or render content in response to the user providing one or more inputs. For example, a user can provide a spoken utterance, such as "Assistant, send a video message to Megan" ("Assistant, send a video message to Megan"). The mobile computing device can generate audio data based on the spoken utterance, and cause the audio data to be processed to identify one or more actions (e.g., initialize a video call to a contact) that the user is requesting.
[0050] When spoken utterances have not yet been detected at the mobile computing device, one or more microphones of the mobile computing device can be monitored for spoken input. However, when spoken utterances are received, method 600 can proceed from operation 602 to operation 604. Operation 604 can include determining whether the requested action involves rendering graphical content. The graphical content can be, but is not limited to, media provided by an application, streaming data, video recorded by a camera accessible to the user, and / or any other video data that may or may not be associated with corresponding audio data. For example, when a user requests to provide a video message to another person, the mobile computing device can determine that the requested action does involve rendering graphical content, because generating a video message can involve rendering a video preview of the video message and rendering a video stream of the recipient (e.g., "Megan").
[0051] When the requested action is determined to involve rendering graphical content, method 600 can proceed from operation 604 to operation 608. Operation 608 can include determining whether the user is within or at a distance for perceiving graphical content. That is, the operation can determine whether the position of the user relative to the mobile computing device satisfies a distance condition. The distance condition can be predetermined, and can be fixed, for example, for all graphical content. Alternatively, the distance condition can vary depending on the specific graphical content (i.e., the distance condition can be determined based on the graphical content). For example, the display of basic content that can be displayed in a large font can be associated with different distance conditions than displaying detailed or densely presented content. In other words, the mobile computing device and / or a server communicating with the mobile computing device can process data to determine whether the user can perceive the graphical content to be displayed at the mobile computing device. For example, the mobile computing device can include a camera that captures image data that can characterize the position of the user relative to the mobile computing device. The mobile computing device can use the image data to determine the proximity of the user relative to the mobile computing device, and thereby determine whether the user can reasonably see the display panel of the mobile computing device. When the mobile computing device determines that the user is not within the distance for perceiving the graphical content (ie, the distance condition associated with the content is satisfied), method 600 can proceed to operation 610 .
[0052] Operation 610 infers maneuvering the mobile computing device within a distance for perceiving the graphical content. In other words, the mobile computing device can operate one or more motors to navigate the mobile computing device toward the user, at least until the mobile computing device reaches or enters a distance for the user to perceive the graphical content. When it is determined that the user is within a distance for perceiving (e.g., able to see and / or read) the graphical content, method 600 can proceed from operation 608 to operation 612.
[0053] When it is determined that the requested action does not involve rendering graphics content, method 600 can proceed to operation 606 from operation 604. Operation 606 can include determining whether the requested action involves rendering audio content. The audio content can include any output from the mobile computing device and / or any other computing device that can be audible to one or more users. When the requested action is determined to involve rendering audio content, method 600 can proceed to operation 616 from the operation at 606. Otherwise, when the requested action is determined not to involve rendering audio content and / or graphics content, method 600 can proceed to operation 614, where the action of one or more requests is initialized in response to a spoken utterance.
[0054] Operation 616 can include determining whether the user is within the distance for perceiving the audio content. In other words, the mobile computing device can determine whether the user's current position will allow the user to hear the audio generated at the mobile computing device or at another computing device that can render the audio content in response to the spoken word. For example, the mobile computing device can generate audio data and / or image data, and the position of the user relative to the mobile computing device can be estimated based on the audio data and / or image data. When the user is not within the distance for perceiving the audio content from the estimated distance of the mobile computing device, method 600 can proceed to operation 618 from operation 616.
[0055] Operation 618 can include causing the mobile computing device to operate within a distance for perceiving audio content. Alternatively or in addition, operation 618 can include determining whether one or more other computing devices are within a distance from the user for rendering audio content. Therefore, if another computing device is located within a distance for rendering audible audio content for the user, the determination at operation 616 can be met with certainty, and method 600 can proceed to operation 620. Otherwise, the mobile computing device can be operated closer to the user so that the mobile computing device will be within a distance for perceiving audio content generated by the mobile computing device for the user. When the mobile computing device is within a distance for receiving audio content for the user, method 600 can proceed to operation 620 from operation 616.
[0056] In the case where the requested action involves graphical content and the mobile computing device has been manipulated to within a distance where the user perceives the graphical content, method 600 can proceed from operation 608 to operation 612. Operation 612 can include causing the mobile computing device to manipulate the display panel to be pointed at the user. The display panel can be controlled by one or more motors attached to one or more housing packages of the mobile computing device. For example, one or more motors can be attached to a first housing package and can be operated to adjust the angle of the display panel. Image data and / or audio data captured at the mobile computing device and / or at any other computing device with permission from the user can be processed to identify one or more anatomical features of the user, such as the user's eyes. Based on the identification of the anatomical features, one or more motors that control the angle of the display panel can be operated to manipulate the display panel so that the display panel projects the graphical content toward the anatomical features of the user. In some embodiments, one or more other motors of the mobile computing device can further adjust the height of the display panel of the mobile computing device. Therefore, one or more motors and / or one or more other motors can be simultaneously operated to manipulate the display panel to be within the user's field of view and / or to point to the user's anatomical features.
[0057] When the mobile computing device has completed manipulating the display panel to point to the user, the method 600 can proceed from the operation at 612 to operation 620. Operation 620 can include causing the requested content to be rendered and / or causing the requested action to be performed. For example, when the user provides a verbal utterance requesting the automated assistant to turn on the lights in the house, the action can involve controlling the IoT device without rendering the audio content and / or display content, thereby allowing the mobile computing device to bypass the direction manipulation toward the user. However, when the verbal utterance includes a request for an audio stream and / or video stream to be provided via the mobile computing device, the mobile computing device can manipulate toward the user and / or confirm that the user is within the distance for perceiving the content. Thereafter, the mobile computing device can then render the content for the user. In this way, delays may otherwise be caused by causing the user to first request the mobile computing device to navigate to the user before rendering the content. In addition, the mobile computing device can save computing resources by choosing whether to navigate to the user depending on the type of content to be rendered for the user. If the mobile computing device navigates to the user without distinction regardless of the (multiple) actions being requested, computing resources such as power and processing bandwidth may be wasted.
[0058] Figure 7 7 is a block diagram of an example computer system 710. The computer system 710 typically includes at least one processor 714 that communicates with a number of peripheral devices via a bus subsystem 712. These peripheral devices may include a storage subsystem 724 including, for example, a memory 725 and a file storage subsystem 726, a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices allow a user to interact with the computer system 710. The network interface subsystem 716 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.
[0059] The user interface input devices 722 may include a keyboard, a pointing device such as a mouse, a trackball, a touch pad or a graphics tablet, a scanner, a touch screen incorporated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into the computer system 710 or over a communication network.
[0060] User interface output device 720 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanisms for creating a visual image. The display subsystem may also provide a non-visual display such as via an audio output device. Typically, the use of the term "output device" is intended to include devices and modes of all possible types of information output from computer system 710 to a user or another machine or computer system.
[0061] The storage subsystem 724 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 724 may include logic for performing selected aspects of the method 600 and / or for implementing one or more of the system 500, the mobile computing device 102, the mobile computing device 204, the mobile computing device 304, the mobile computing device 404, the automated assistant, the computing device 518, the server device 502, and / or any other application, device, apparatus, and / or module discussed herein.
[0062] These software modules are usually executed by the processor 714 alone or in combination with other processors. The memory 725 used in the storage subsystem 724 can include multiple memories, including a main random access memory (RAM) 730 for storing instructions and data during program execution and a read-only memory (ROM) 732 in which fixed instructions are stored. The file storage subsystem 726 can provide permanent storage for program and data files, and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media box. The modules that implement the functions of certain embodiments can be stored in the storage subsystem 724 by the file storage subsystem 726, or stored in other machines accessible by the processor 714.
[0063] The bus subsystem 712 provides a mechanism for the various components and subsystems of the computer system 710 to communicate with each other as desired. Although the bus subsystem 712 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.
[0064] Computer system 710 can be of various types including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 7 The description of the computer system 710 depicted in FIG. 7 is intended only as a specific example for the purpose of illustrating some embodiments. Many other configurations of the computer system 710 may have more Figure 7The computer system depicted in FIG.
[0065] In cases where the systems described herein collect personal information about users (or, as they are often referred to herein, "participants") or may utilize personal information, users may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, the user's preferences, or the user's current geographic location) or to control whether and / or how content that may be more relevant to the user is received from a content server. In addition, certain data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, the user's identity may be processed so that personally identifiable information cannot be determined for the user, or the user's geographic location may be summarized as a place where geographic location information is obtained (such as to a city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, users may control how information about users is collected and / or used.
[0066] Although several embodiments have been described and illustrated herein, various other means and / or structures for performing the functions described herein and / or obtaining the results and / or one or more advantages described herein may be utilized, and each of such variations and / or modifications is considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications in which the present teachings are used.
[0067] Those skilled in the art will recognize or be able to ascertain many equivalents of the specific embodiments described herein using no more than routine experiments. Therefore, it should be understood that the aforementioned embodiments are presented by way of example only, and within the scope of the appended claims and their equivalents, the embodiments may be practiced in a manner different from that specifically described and claimed. Embodiments of the present disclosure relate to each individual feature, system, article, material, kit and / or method described herein. In addition, if these features, systems, articles, materials, kits and / or methods are not mutually contradictory, any combination of two or more of these features, systems, articles, materials, kits and / or methods is included within the scope of the present disclosure.
[0068] In some embodiments, the method is described as including an operation such as determining that a user has provided a spoken utterance based on input from one or more microphones of a mobile computing device, wherein the mobile computing device includes one or more first motors that manipulate the mobile computing device across an area. The method can further include determining that the user is requesting the mobile computing device to perform an action associated with rendering content via one or more speakers and / or a display panel of the automated assistant via the mobile computing device based on input from the one or more microphones, additional input from one or more microphones and / or one or more other sensors of the mobile computing device, determining the user's position relative to the mobile computing device. The method can further include, when the content requested by the user to be rendered at the mobile computing device includes graphical content and when the determined position satisfies a specific distance condition: causing the first motor of the mobile computing device to manipulate the mobile computing device toward the user's position, and causing the display panel to render the graphical content to facilitate performing the action.
[0069] In some implementations, the method can further include, when the content requested by the user to be rendered at the mobile computing device includes audio content: determining whether the mobile computing device is within a distance from the user for audibly rendering the audio content for the user. The method can further include, when the mobile computing device is not within a distance from the user for audibly rendering the audio content for the user: based on determining that the mobile computing device is not within the distance from the user, causing one or more first motors of the mobile computing device to manipulate the mobile computing device toward the position of the user, and causing one or more speakers of the mobile computing device to render the audio content to facilitate performing the action.
[0070] In some embodiments, the method can further include, when the content requested by the user to be rendered at the mobile computing device includes graphical content and when the determined location satisfies the distance condition: causing one or more second motors of the mobile computing device to manipulate a display panel of the mobile computing device to facilitate rendering of the graphical content toward the user. In some embodiments, the method can further include determining whether a user is providing a subsequent spoken utterance while the one or more first motors and / or the one or more second motors of the mobile computing device are operating; and when the subsequent spoken utterance is received while the one or more first motors and / or the one or more second motors of the mobile computing device are operating: causing the one or more first motors and / or the one or more second motors to transition to a reduced power state, wherein the reduced power state corresponds to a state in which the one or more first motors and / or the one or more second motors consume less power than another state and / or a previous state of the one or more first motors and / or the one or more second motors.
[0071] In some embodiments, the method can further include, when the subsequent spoken utterance is no longer received while the one or more first motors and / or the one or more second motors of the mobile computing device are operating: causing the one or more first motors and / or the one or more second motors to transition from a reduced power state to another operating state to facilitate manipulating the display panel and / or manipulating the mobile computing device toward the location of the user. In some embodiments, the method can further include, in response to receiving input to a microphone and using a camera of the mobile computing device to identify an anatomical feature of the user, wherein causing the one or more second motors to manipulate the display panel includes pointing the display panel toward the anatomical feature of the user. In some embodiments, the method can further include, when the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: causing one or more third motors of the mobile computing device to manipulate the camera of the mobile computing device to be directed toward the user based on the content requested by the user corresponding to graphical content and / or audio content.
[0072] In some embodiments, the display panel is mounted to a first housing package of the mobile computing device, the camera is mounted to a second housing package of the mobile computing device, and one or more third motors are at least partially encapsulated within the third housing package of the mobile computing device. In some embodiments, causing the one or more third motors of the mobile computing device to manipulate the camera of the mobile computing device in the direction of the user includes: rotating the second housing package of the mobile computing device around an axis that intersects the third housing package of the mobile computing device. In some embodiments, the fourth motor is at least partially encapsulated in the second housing package and controls radial movement of the second housing package relative to the third housing package, and the method further includes: when the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: causing the fourth motor to achieve radial movement of the second housing package so that the second housing package changes a separation angle from the third housing package.
[0073] In some embodiments, the method can further include, when the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: in response to receiving input to the microphone and using a camera of the mobile computing device, identifying an anatomical feature of the user, and based on identifying the anatomical feature of the user, determining a separation angle of the second shell package relative to the third shell package, wherein the separation angle corresponds to the angle at which the camera is pointed at the anatomical feature of the user. In some embodiments, a fifth motor is at least partially encapsulated in the first shell package and controls another radial movement of the first shell package relative to the second shell package, and the method further includes: when the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: causing the fifth motor to effect another radial movement of the first shell package so that the first shell package reaches another separation angle from the second shell package.
[0074] In some embodiments, the method can further include, when the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: in response to receiving input to the microphone and using a camera of the mobile computing device, identifying an anatomical feature of the user, and based on identifying the anatomical feature of the user, determining another angle of separation of the first housing enclosure relative to the second housing enclosure, wherein the other angle of separation corresponds to another angle at which the display panel is directed toward the anatomical feature of the user. In some embodiments, determining the position of the user relative to the mobile computing device includes: using outputs from a plurality of microphones of the mobile computing device to determine that the position includes a plurality of different persons, and using other outputs from the camera of the mobile computing device to determine that the user is one of the plurality of different persons. In some embodiments, the method can further include, after rendering the display content and / or audio content, causing one or more second motors to lower the height of the mobile computing device by manipulating the first housing enclosure and the display panel of the mobile computing device toward the second housing enclosure of the mobile computing device.
[0075] In other embodiments, the method is described as including operations such as determining that a user has provided a spoken utterance to the mobile computing device based on input to one or more microphones of the mobile computing device. The method can further include, in response to the spoken utterance being provided to the mobile computing device, causing one or more motors of the mobile computing device to manipulate a display panel attached to a first housing package of the mobile computing device away from a second housing package of the mobile computing device. The method can further include determining whether another spoken utterance is directed at the mobile computing device while the one or more motors are manipulating the first housing package away from the second housing package. The method can further include, when it is determined that another spoken utterance is directed at the mobile computing device: causing one or more motors to transition to a low power state while the other spoken utterance is directed at the mobile computing device, causing an automated assistant accessible via the mobile computing device to initialize the execution of an action based on the other spoken utterance. The method can further include, when the other spoken utterance is completed and / or no longer directed at the mobile computing device: causing one or more motors to complete the manipulation of the first housing package away from the second housing package.
[0076] In some implementations, the method can further include, in response to the spoken utterance being provided to the mobile computing device, causing one or more second motors of the mobile computing device to drive the mobile computing device toward the location of the user. In some implementations, the method can further include, when it is determined that another spoken utterance is directed toward the mobile computing device: causing the one or more second motors of the mobile computing device to pause driving the mobile computing device toward the location of the user, and when the another spoken utterance is completed and / or is no longer directed toward the mobile computing device: causing the one or more second motors of the mobile computing device to continue driving the mobile computing device toward the location of the user.
[0077] In some embodiments, the method can further include, when the one or more second motors have completed driving the mobile computing device toward the user's location: based on the spoken utterance and / or another spoken utterance, causing one or more third motors of the mobile computing device to manipulate the second housing enclosure away from the third housing enclosure of the mobile computing device and manipulate the camera of the mobile computing device toward the user. In some embodiments, the method can also include, when the one or more second motors have completed driving the mobile computing device toward the user's location: based on the spoken utterance and / or another spoken utterance, causing one or more fourth motors to rotate the first housing enclosure about an axis that intersects a surface of the third housing enclosure to facilitate pointing the display panel toward the user.
[0078] In other embodiments, the method is described as including operations such as determining that a user has provided a spoken utterance to a mobile computing device based on input from one or more microphones of the mobile computing device. The method can further include, in response to the spoken utterance being provided to the mobile computing device, causing one or more motors of the mobile computing device to manipulate the mobile computing device toward the user's location. The method can further include determining whether another spoken utterance is directed at the mobile computing device while the one or more motors are manipulating the mobile computing device toward the user's location. The method can further include, when it is determined that the other spoken utterance is directed at the mobile computing device: causing the one or more motors to transition to a low power state while the other spoken utterance is directed at the mobile computing device, so that an automated assistant accessible via the mobile computing device initializes the performance of an action based on the other spoken utterance. The method can further include, when the other spoken utterance is completed and / or is no longer directed at the mobile computing device: causing the one or more motors to continue manipulating the mobile computing device toward the user's location.
[0079] In some embodiments, the method can further include, in response to a spoken utterance being provided to the mobile computing device, causing one or more second motors of the mobile computing device to manipulate a display panel attached to a first housing package of the mobile computing device away from a second housing package of the mobile computing device. In some embodiments, the method can further include, when it is determined that another spoken utterance is directed to the mobile computing device: causing one or more second motors to transition to a low power state while the other spoken utterance is directed to the mobile computing device, causing an automated assistant accessible via the mobile computing device to initiate performance of an action based on the other spoken utterance. In some embodiments, the method can further include, when the other spoken utterance is completed and / or no longer directed to the mobile computing device: causing the one or more motors to complete manipulating the first housing package away from the second housing package.
[0080] In some embodiments, the method can further include, when the one or more motors have completed manipulating the mobile computing device toward the user's position: based on the spoken utterance and / or another spoken utterance, causing one or more third motors of the mobile computing device to manipulate the second housing enclosure away from the third housing enclosure of the mobile computing device and manipulate the camera of the mobile computing device toward the user. In some embodiments, the method can further include, when the one or more motors have completed manipulating the mobile computing device toward the user's position: based on the spoken utterance and / or another spoken utterance, causing one or more fourth motors to rotate the first housing enclosure about an axis that intersects a surface of the third housing enclosure of the mobile computing device to facilitate directing the display panel toward the user.
Claims
1. A method for rendering content to a user, include: determining, based on input to one or more microphones of a mobile computing device, that a user has provided a spoken utterance, wherein the mobile computing device includes one or more first motors that manipulate the mobile computing device across an area; determining, based on the input to the one or more microphones, that the user is requesting the mobile computing device to perform an action associated with an automated assistant rendering content via one or more speakers and / or display panels of the mobile computing device; determining a position of the user relative to the mobile computing device based on the input to the one or more microphones, additional input to the one or more microphones and / or one or more other sensors of the mobile computing device; and When the content requested by the user to be rendered at the mobile computing device includes graphical content and the user is not located within a distance for perceiving the graphical content: causing the first motor of the mobile computing device to maneuver the mobile computing device toward the location of the user, subsequently, when the user is located within the distance for perceiving the graphical content, causing one or more second motors of the mobile computing device to manipulate the display panel of the mobile computing device to facilitate rendering the graphical content to the user, and The display panel is caused to render the graphical content to facilitate performing the action.
2. The method according to claim 1, further comprising: include: When the content requested by the user to be rendered at the mobile computing device includes audio content: determining whether the mobile computing device is within a distance from the user for audibly rendering the audio content to the user, and When the mobile computing device is not within the distance from the user for audibly rendering the audio content for the user: Based on determining that the mobile computing device is not within the distance from the user, causing the one or more first motors of the mobile computing device to maneuver the mobile computing device toward the location of the user, and One or more speakers of the mobile computing device are caused to render the audio content to facilitate performing the action.
3. The method according to claim 1, further comprising: include: determining whether the user is providing a subsequent spoken utterance while the one or more first motors and / or the one or more second motors of the mobile computing device are operating; as well as When the subsequent spoken utterance is received while the one or more first motors and / or the one or more second motors of the mobile computing device are operating: Converting the one or more first motors and / or the one or more second motors to a reduced power state, wherein the reduced power state corresponds to a state in which the one or more first motors and / or the one or more second motors consume less power than another state and / or a previous state of the one or more first motors and / or the one or more second motors.
4. The method according to claim 3, further comprising: include: When the subsequent spoken utterance is no longer received while the one or more first motors and / or the one or more second motors of the mobile computing device are operating: The one or more first motors and / or the one or more second motors are transitioned from the reduced power state to the another state to facilitate manipulating the display panel and / or manipulating the mobile computing device toward the location of the user.
5. The method according to claim 1, further comprising: include: In response to receiving the input to the microphone and identifying an anatomical feature of the user using a camera of the mobile computing device, wherein causing the one or more second motors to manipulate the display panel includes pointing the display panel toward the anatomical feature of the user.
6. The method according to claim 1, further comprising: include: When the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: Based on the content requested by the user corresponding to graphical content and / or audio content, one or more third motors of the mobile computing device are caused to manipulate a camera of the mobile computing device to be directed toward the user.
7. The method according to claim 6, in, The display panel is mounted to a first housing enclosure of the mobile computing device, the camera is mounted to a second housing enclosure of the mobile computing device, and the one or more third motors are at least partially enclosed within a third housing enclosure of the mobile computing device.
8. The method according to claim 7, in, Causing the one or more third motors of the mobile computing device to steer the camera of the mobile computing device in the direction of the user includes: The second housing enclosure of the mobile computing device is rotated about an axis that intersects the third housing enclosure of the mobile computing device.
9. The method according to claim 7, in, A fourth electric machine is at least partially encapsulated in the second housing enclosure and controls radial movement of the second housing enclosure relative to the third housing enclosure, and the method further comprises: When the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: The fourth motor is enabled to realize the radial movement of the second housing packaging, so that the second housing packaging changes a separation angle from the third housing packaging.
10. The method according to claim 9, further comprising: include: When the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: In response to receiving the input to the microphone and using the camera of the mobile computing device, identifying an anatomical feature of the user, and Based on identifying the anatomical feature of the user, the separation angle of the second housing enclosure relative to the third housing enclosure is determined, wherein the separation angle corresponds to an angle at which the camera is pointed toward the anatomical feature of the user.
11. The method according to claim 7, in, A fifth electric machine is at least partially enclosed in the first housing enclosure and controls another radial movement of the first housing enclosure relative to the second housing enclosure, and the method further comprises: When the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: The fifth motor is caused to realize another radial movement of the first housing enclosure, so that the first housing enclosure reaches another separation angle from the second housing enclosure.
12. The method according to claim 11, further comprising: include: When the content requested by the user to be rendered at the mobile computing device corresponds to graphical content and / or audio content: In response to receiving the input to the microphone and using the camera of the mobile computing device, identifying an anatomical feature of the user, and Based on identifying the anatomical feature of the user, the other separation angle of the first shell enclosure relative to the second shell enclosure is determined, wherein the other separation angle corresponds to another angle at which the display panel points toward the anatomical feature of the user.
13. The method according to claim 1, in, Determining the location of the user relative to the mobile computing device includes: using outputs from a plurality of microphones of the mobile computing device to determine that the location includes a plurality of different persons, and Other output from the camera of the mobile computing device is used to determine that the user is one of the plurality of different persons.
14. The method according to claim 2, further comprising: include: After rendering the graphical content and / or the audio content, causing the one or more second motors to lower the height of the mobile computing device by manipulating the first body enclosure of the mobile computing device and the display panel toward the second body enclosure of the mobile computing device.
15. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-14.
16. A computer-readable storage medium comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 14.
17. A system for rendering content to a user, comprising one or more processors for performing the method of any one of claims 1 to 14.
Citation Information
Patent Citations
Display control device, display system, display device, terminal device, display control method and program
JP2014013494A
Communication robot system
JP2018067785A
Robot
US20100180709A1