Detection of Visual Attention during User Speech
By using simultaneous audio and video streams to determine user attention, the electronic device accurately initiates tasks without additional inputs, improving interaction efficiency and reducing power consumption.
Patent Information
- Application Number
- JP2024569639
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-10
- Filing Date
- 2023-05-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-05-22
AI Technical Summary
Existing digital assistants struggle to accurately determine whether a user's visual attention is directed towards the device while speaking, leading to inefficient and inaccurate responses, as they often require additional user inputs to confirm the intended target of speech commands.
An electronic device simultaneously receives audio and video streams to determine the user's visual attention, identifies the portion of the audio stream containing a user utterance addressed to the device, and initiates a task based on this determination, eliminating the need for explicit speech trigger inputs.
This method enhances the accuracy and efficiency of user-device interaction by reducing the need for additional user inputs, preventing erroneous responses, and conserving power by ensuring quicker and more relevant device operations.
Smart Images

Figure 2025520085000001_ABST
Abstract
Description
Cross - reference to related applications
[0001] This application claims priority to U.S. Patent Application No. 18 / 132,793, entitled "DETECTING VISUAL ATTENTION DURING USER SPEECH", filed on April 10, 2023, U.S. Provisional Patent Application No. 63 / 456,639, entitled "DETECTING VISUAL ATTENTION DURING USER SPEECH", filed on April 3, 2023, and U.S. Provisional Patent Application No. 63 / 346,693, entitled "DETECTING VISUAL ATTENTION DURING USER SPEECH", filed on May 27, 2022. The entire contents of each of these applications are hereby incorporated by reference into this specification.
Technical Field
[0002] This generally relates to determining whether a user's visual attention is directed towards an electronic device while the user is speaking.
Background Art
[0003] Intelligent automatic assistants (or digital assistants) can provide a useful interface between human users and electronic devices. With such assistants, users may be able to interact with a device or system using natural language in oral and / or text form. For example, a user can provide an utterance input containing a user request to a digital assistant operating on an electronic device. The digital assistant can interpret the user's intention from the utterance input and make the user's intention operable for a task. The task can then be executed by performing one or more services of the electronic device and return a relevant output response to the user request to the user.
Summary of the Invention
[0004] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors and memory, receiving an audio stream and a video stream simultaneously; determining, based on a first portion of the audio stream received within a predetermined duration prior to the current time and a first portion of the video stream received within a predetermined duration prior to the current time, whether the user's visual attention is directed to the electronic device while the user is speaking; identifying a second portion of the audio stream to include a user utterance addressed to the electronic device in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking; starting a task based on the second portion of the audio stream by a digital assistant operating on the electronic device; and providing an output indicative of the started task.
[0005] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of the electronic device, cause the electronic device to receive an audio stream and a video stream simultaneously; determine, based on a first portion of the audio stream received within a predetermined duration prior to the current time and a first portion of the video stream received within a predetermined duration prior to the current time, whether the user's visual attention is directed to the electronic device while the user is speaking; identify a second portion of the audio stream to include a user utterance addressed to the electronic device in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking; start a task based on the second portion of the audio stream by a digital assistant operating on the electronic device; and provide an output indicative of the started task.
[0006] Exemplary electronic devices are disclosed herein. The exemplary electronic device includes one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions that receive an audio stream and a video stream simultaneously, and based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within a predetermined duration before the current time, determine whether the user's visual attention is directed to the electronic device while the user is speaking, and while the user is speaking, in accordance with the determination that the user's visual attention is directed to the electronic device, identify a second portion of the audio stream that includes the user utterance addressed to the electronic device, and by a digital assistant operating on the electronic device, based on the second portion of the audio stream, start a task and provide an output indicating the started task.
[0007] The exemplary electronic device includes means for receiving an audio stream and a video stream simultaneously, determining whether the user's visual attention is directed to the electronic device while the user is speaking based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within a predetermined duration before the current time, identifying a second portion of the audio stream that includes the user utterance addressed to the electronic device in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking, and by a digital assistant operating on the electronic device, based on the second portion of the audio stream, starting a task and providing an output indicating the started task.
[0008] Determining that the user's visual attention is directed to the electronic device while the user is speaking and identifying a second portion of the audio stream can enable the device to respond more accurately and efficiently to the speech input. For example, the user can simply look at the device (or a part thereof) while speaking a user request to the device to provide a relevant response. Further, to provide a relevant response, the device may not need additional user input that explicitly indicates that the speech input is directed to the device, such as a speech trigger input, a button selection, a selection of a displayed affordance, etc. In this way, the user-device interface can be more efficient and accurate (e.g., by reducing the amount of user input required to operate the device, by reducing the user input required to prevent the device from erroneously responding to speech input not directed to the device, by accurately responding to speech input directed to the device, by avoiding repeated speech input to the device), which in turn can reduce power usage and improve the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0009] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors and memory, simultaneously receiving an audio stream and a video stream; identifying a first portion of the video stream to include a first user; identifying a second portion of the video stream to include a second user; identifying a first portion of the audio stream that includes an utterance of the first user addressed to the electronic device, according to a determination that a visual attention of the first user is directed to the electronic device while the first user is speaking, based on the audio stream and the first portion of the video stream; providing a first output based on processing the first portion of the audio stream; identifying a second portion of the audio stream that includes an utterance of the second user addressed to the electronic device, according to a determination that a visual attention of the second user is directed to the electronic device while the second user is speaking, based on the audio stream and the second portion of the video stream; and providing a second output based on processing the second portion of the audio stream.
[0010] Exemplary non-transitory computer-readable media are disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive an audio stream and a video stream simultaneously, identify a first portion of the video stream that includes a first user, identify a second portion of the video stream that includes a second user, based on the audio stream and the first portion of the video stream, in accordance with a determination that the visual attention of the first user is directed to the electronic device while the first user is speaking, identify a first portion of the audio stream that includes the utterance of the first user addressed to the electronic device, provide a first output based on processing the first portion of the audio stream, based on the audio stream and the second portion of the video stream, in accordance with a determination that the visual attention of the second user is directed to the electronic device while the second user is speaking, identify a second portion of the audio stream that includes the utterance of the second user addressed to the electronic device, and provide a second output based on processing the second portion of the audio stream.
[0011] Exemplary electronic devices are disclosed herein. The exemplary electronic device includes one or more processors, a memory, and one or more programs. The one or more programs are stored in the memory and configured to be executed by the one or more processors. The one or more programs include instructions that receive an audio stream and a video stream simultaneously, identify a first portion of the video stream that includes a first user, identify a second portion of the video stream that includes a second user, identify a first portion of the audio stream that includes utterances of the first user addressed to the electronic device according to a determination that the visual attention of the first user is directed to the electronic device while the first user is speaking, based on the audio stream and the first portion of the video stream, provide a first output based on processing the first portion of the audio stream, identify a second portion of the audio stream that includes utterances of the second user addressed to the electronic device according to a determination that the visual attention of the second user is directed to the electronic device while the second user is speaking, based on the audio stream and the second portion of the video stream, and provide a second output based on processing the second portion of the audio stream.
[0012] An exemplary electronic device receives an audio stream and a video stream simultaneously, identifies a first portion of the video stream that includes a first user, identifies a second portion of the video stream that includes a second user, based on the audio stream and the first portion of the video stream, in accordance with a determination that the visual attention of the first user is directed to the electronic device while the first user is speaking, identifies a first portion of the audio stream that includes speech of the first user addressed to the electronic device, provides a first output based on processing the first portion of the audio stream, based on the audio stream and the second portion of the video stream, in accordance with a determination that the visual attention of the second user is directed to the electronic device while the second user is speaking, identifies a second portion of the audio stream that includes speech of the second user addressed to the electronic device, and includes means for providing a second output based on processing the second portion of the audio stream.
[0013] Providing an output based on processing an identified portion of an audio stream as certain conditions are met can enable a device to accurately and efficiently respond to the speech of the correct user in a multi-user environment. For example, in a multi-user environment, to respond to a user's speech, the device determines that the same user is looking at the device while speaking. Thus, the device can avoid incorrectly responding to the speech of a user not addressed to the device. For example, if a first user is looking at the device while not speaking and a second user speaks while not looking at the device, the device does not incorrectly respond to the second user's speech. In this way, the user-device interface can be more efficient and accurate (e.g., by reducing the amount of user input required to operate the device, by reducing the user input required to stop the device from responding to the speech of the wrong user, by accurately identifying and responding to speech input from the correct user, by avoiding repeated speech input to the device), which further reduces power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0014] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors and a memory, receiving an audio stream and a video stream simultaneously, determining a first type of reliability score indicative of whether the user's visual attention is directed to the electronic device while the user is speaking, based on the audio stream and the video stream, after determining the first type of reliability score, determining a final reliability score indicative of whether the user's visual attention is directed to the electronic device while the user is speaking, based on the first type of reliability score and a second type of reliability score indicative of whether the user's visual attention is directed to the electronic device, identifying a portion of the audio stream that includes the user utterance addressed to the electronic device according to a determination that the final reliability score exceeds a threshold, and providing an output based on processing the portion of the audio stream.
[0015] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of the electronic device, cause the electronic device to receive an audio stream and a video stream simultaneously, determine a first type of reliability score indicative of whether the user's visual attention is directed to the electronic device while the user is speaking, based on the audio stream and the video stream, after determining the first type of reliability score, determine a final reliability score indicative of whether the user's visual attention is directed to the electronic device while the user is speaking, based on the first type of reliability score and a second type of reliability score indicative of whether the user's visual attention is directed to the electronic device, identify a portion of the audio stream that includes the user utterance addressed to the electronic device according to a determination that the final reliability score exceeds a threshold, and provide an output based on processing the portion of the audio stream.
[0016] Exemplary electronic devices are disclosed herein. The exemplary electronic device includes one or more processors, a memory, and one or more programs. The one or more programs are stored in the memory and configured to be executed by the one or more processors. The one or more programs include instructions that receive an audio stream and a video stream simultaneously, and based on the audio stream and the video stream, determine a first type of reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking. After determining the first type of reliability score, based on the first type of reliability score and a second type of reliability score indicating whether the user's visual attention is directed to the electronic device, determine a final reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking. According to the determination that the final reliability score exceeds a threshold, identify a portion of the audio stream that includes the user utterance addressed to the electronic device, and provide an output based on processing the portion of the audio stream.
[0017] The exemplary electronic device includes means for receiving an audio stream and a video stream simultaneously, and based on the audio stream and the video stream, determining a first type of reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking. After determining the first type of reliability score, based on the first type of reliability score and a second type of reliability score indicating whether the user's visual attention is directed to the electronic device, determining a final reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking. According to the determination that the final reliability score exceeds a threshold, identifying a portion of the audio stream that includes the user utterance addressed to the electronic device, and providing an output based on processing the portion of the audio stream.
[0018] Identifying a portion of an audio stream according to a determination that a final reliability score exceeds a threshold can enable a device to respond more accurately and efficiently to an utterance input. For example, determining the final reliability score can enable the device to consider additional relevant factor(s) (e.g., the user's posture, the user's line of sight, relative motion between the user and the device, etc.) in order to more accurately determine that the user is looking at the device while speaking. As explained, determining whether the user is looking at the device while speaking can enable the device to respond efficiently and accurately to the utterance input, for example, without requiring additional user input. In this way, the user-device interface can be more efficient and accurate (e.g., by reducing the amount of user input required to operate the device, by reducing the user input required to prevent the device from erroneously responding to an utterance input not directed to the device, by accurately responding to an utterance input directed to the device, by avoiding repeated utterance inputs to the device), which further reduces power usage and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.
Brief Description of the Drawings
[0019]
Figure 1
[0020]
Figure 2A
[0021]
Figure 2B
[0022]
Figure 3
[0023]
Figure 4
[0024]
Figure 5A
[0025]
Figure 5B
[0026]
Figure 6A
[0027]
Figure 6B
[0028]
Figure 7A
[0029]
Figure 7B
[0030]
Figure 7C
[0031]
Figure 8A
[0032]
Figure 8B
[0033]
Figure 9
[0034]
Figure 10A
Figure 10B
[0035]
Figure 11
[0036]
Figure 12
[0037]
Figure 13
DETAILED DESCRIPTION OF THE INVENTION
[0038] In the following description of the embodiments, reference is made to the accompanying drawings, which show specific embodiments that are practical as examples. It should be understood that other embodiments can be used and structural changes can be made without departing from the scope of the various embodiments.
[0039] This generally relates to determining whether a user's visual attention is directed towards an electronic device while the user is speaking. If the user's visual attention is directed towards the electronic device while the user is speaking, the electronic device can identify a portion of the received audio stream to include utterances addressed to the electronic device. The electronic device can then process the identified portion, for example, using a digital assistant, and provide a relevant output to the user.
[0040] In the following description, terms such as "first", "second", etc. are used to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various embodiments described, the first input can be referred to as the second input, and similarly, the second input can be referred to as the first input. The first input and the second input are both inputs, and in some cases, distinct and different inputs.
[0041] The terminology used in the description of the various embodiments described herein is for the purpose of describing particular embodiments only and is not intended to be limiting. When used in the description of the various embodiments described and in the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Also, as used herein, the term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items. It is to be understood that the terms "includes", "including", "comprises", and / or "comprising", when used herein, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0042] The term "if (in the case of ~)" can be interpreted, depending on the context, to mean "when (when ~)", or "upon (~ then)", or "in response to determining (in response to the determination that ~)", or "in response to detecting (in response to the detection of ~)". Similarly, the phrase "if it is determined (if it is determined that ~)" or "if [a stated condition or event] is detected (if [a stated condition or event] is detected)" can be interpreted, depending on the context, to mean "upon determining (when determining that ~)", or "in response to determining (in response to the determination that ~)", or "upon detecting [the stated condition or event] (when detecting [the stated condition or event])", or "in response to detecting [the stated condition or event] (in response to the detection of [the stated condition or event])". 1. Systems and Environments
[0043] FIG. 1 shows a block diagram of system 100 according to various embodiments. In some embodiments, system 100 executes a digital assistant. The terms “digital assistant,” “virtual assistant,” “intelligent automated assistant,” or “automated digital assistant” refer to any information processing system that infers user intent by interpreting natural language input in oral and / or text form and performs an action based on the inferred user intent. For example, to operate based on the inferred user intent, the system performs one or more of the following: identifying a task flow having steps and parameters designed to fulfill the inferred user intent, inputting specific requirements from the inferred user intent into the task flow, executing the task flow by calling a program, method, service, API, or the like, and generating an output response to the user in audible (e.g., speech) and / or visual form.
[0044] Specifically, the digital assistant can receive user requests in the form of at least partially natural language commands, requests, opinions, conversations, and / or inquiries. Typically, a user request seeks either an information answer by the digital assistant or the execution of a task. A satisfactory response to a user request includes providing the requested information answer, executing the requested task, or a combination of the two. For example, a user may ask the digital assistant a question such as "Where am I right now?" Based on the user's current location, the digital assistant responds with "You are in Central Park near the West Gate." A user may also request the execution of a task, such as "Please invite my friends to my girlfriend's birthday party next week." In response, the digital assistant can state "Yes, right away." and affirmatively respond to the request, and then, on behalf of the user, send suitable calendar invitations to each of the user's friends listed in the user's email address book. During the implementation of the requested task, the digital assistant may interact with the user in a continuous conversation involving multiple information exchanges over a long period of time. There are many other ways to interact with the digital assistant to request information or the execution of various tasks. In addition to providing a verbal response and taking programmed actions, the digital assistant can also provide responses in other visual or audio formats, such as text, alerts, music, videos, animations, etc.
[0045] As shown in FIG. 1, in some embodiments, the digital assistant is executed according to a client-server model. The digital assistant includes a client-side portion 102 (hereinafter, “DA client 102”) executed on the user device 104 and a server-side portion 106 (hereinafter, “DA server 106”) executed on the server system 108. The DA client 102 communicates with the DA server 106 through one or more networks 110. The DA client 102 provides client-side functions such as user-responsive input and output processing and communication with the DA server 106. The DA server 106 provides server-side functions to any number of DA clients 102 each resident on an individual user device 104.
[0046] In some embodiments, the DA server 106 includes a client-responsive I / O interface 112, one or more processing modules 114, data and models 116, and an I / O interface 118 to external services. The client-responsive I / O interface 112 facilitates client-responsive input and output processing of the DA server 106. The one or more processing modules 114 utilize the data and models 116 to process the speech input and determine the user's intent based on the natural language input. Further, the one or more processing modules 114 perform task execution based on the inferred user intent. In some embodiments, the DA server 106 communicates with an external service 120 through the network(s) 110 for task completion or information acquisition. The I / O interface 118 to external services facilitates such communication.
[0047] The user device 104 can be any suitable electronic device. In some embodiments, the user device 104 is a portable multifunctional device (e.g., the device 200 described below in connection with FIG. 2A), a multifunctional device (e.g., the device 400 described below in connection with FIG. 4), or a personal electronic device (e.g., the device 600 described below in connection with FIGS. 6A - 6B). The portable multifunctional device is, for example, a cellular phone that also includes other functions such as a PDA and / or music player functions. Specific examples of the portable multifunctional device include the Apple Watch (registered trademark), iPhone (registered trademark), iPod Touch (registered trademark), and iPad (registered trademark) devices by Apple Inc. (Cupertino, California). Other examples of the portable multifunctional device include, without limitation, earphones / headphones, speakers, and laptop computers or tablet computers. Further, in some embodiments, the user device 104 is a non - portable multifunctional device. Specifically, the user device 104 is a desktop computer, a game console, a speaker, a television, or a television set - top box. In some embodiments, the user device 104 includes a touch - sensitive surface (e.g., a touch - screen display and / or a touch - pad). Further, the user device 104 optionally includes one or more other physical user interface devices such as a physical keyboard, a mouse, and / or a joystick. Various embodiments of electronic devices such as multifunctional devices are described in more detail below.
[0048] Examples of the communication network(s) 110 include a local area network (LAN) and a wide area network (WAN), such as the Internet. The communication network(s) 110 is implemented using any known network protocol, including various wired or wireless protocols such as Ethernet, Universal Serial Bus (USB), FIREWIRE, Global System for Mobile Communications (GSM) for mobile communication, Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth (registered trademark), Wi-Fi (registered trademark), voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.
[0049] The server system 108 is implemented on one or more stand-alone data processing devices or on a distributed computer network. In some embodiments, the server system 108 also employs the services of various virtual devices and / or third-party service providers (e.g., third-party cloud service providers) to provide the basic computing resources and / or infrastructure resources of the server system 108.
[0050] In some embodiments, user device 104 communicates with DA server 106 via a second user device 122. The second user device 122 is similar or identical to user device 104. For example, the second user device 122 is similar to device 200, 400, or 600 described below in connection with FIGS. 2A, 4, and 6A - 6B. User device 104 is configured to be communicatively coupled to the second user device 122 via a direct communication connection such as Bluetooth, NFC, BTLE, or via a wired or wireless network such as a local Wi-Fi network. In some embodiments, the second user device 122 is configured to act as a proxy between user device 104 and DA server 106. For example, the DA client 102 of user device 104 is configured to send information (e.g., a user request received at user device 104) to DA server 106 via the second user device 122. DA server 106 processes the information and returns relevant data (e.g., data content in response to the user request) to user device 104 via the second user device 122.
[0051] In some embodiments, the user device 104 is configured to reduce the amount of information transmitted from the user device 104 by communicating with a second user device 122 in response to a shortened request regarding data. The second user device 122 is configured to determine supplementary information to add to the shortened request and generate a complete request for transmission to the DA server 106. This system architecture advantageously allows a user device 104 having limited communication capabilities and / or limited battery power (e.g., a wristwatch or similar small electronic device) to access services provided by the DA server 106 by using a second user device 122 having higher communication capabilities and / or battery power (e.g., a mobile phone, laptop computer, tablet computer, etc.) as a proxy to the DA server 106. Although only two user devices 104 and 122 are shown in FIG. 1, it should be understood that the system 100 may include any number and type of user devices configured to communicate with the DA server system 106 in this proxy configuration in some embodiments.
[0052] The digital assistant shown in FIG. 1 includes both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106). However, in some embodiments, the functionality of the digital assistant is implemented as a stand-alone application installed on the user device. Further, the division of functionality between the client and server portions of the digital assistant may vary depending on the implementation. For example, in some embodiments, the DA client is a thin client that provides only user-facing input and output processing functionality and delegates all other functionality of the digital assistant to a backend server. 2. Electronic Device
[0053] Attention is now directed to an embodiment of an electronic device for executing a client-side portion of a digital assistant. FIG. 2A is a block diagram showing a portable multifunctional device 200 including a touch-sensing display system 212 according to some embodiments. The touch-sensing display 212 may be referred to herein as a "touch screen" for convenience, and may be known or referred to as a "touch-sensing display system". Device 200 includes a memory 202 (optionally including one or more computer-readable storage media), a memory controller 222, one or more processing units (CPUs) 220, a peripheral device interface 218, an RF circuit 208, an audio circuit 210, a speaker 211, a microphone 213, an input / output (I / O) subsystem 206, other input control devices 216, and an external port 224. Device 200 optionally includes one or more optical sensors 264. Device 200 optionally includes one or more contact intensity sensors 265 (e.g., a touch-sensing surface such as the touch-sensing display system 212 of device 200) for detecting the intensity of contact on device 200. Device 200 optionally includes one or more haptic output generators 267 for generating haptic output on device 200 (e.g., on a touch-sensing surface such as the touch-sensing display system 212 of device 200 or the touch pad 455 of device 400). These components optionally communicate via one or more communication buses or signal lines 203.
[0054] As used herein and in the claims, the term "intensity" of a contact on a touch sensing surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch sensing surface, or a proxy for the force or pressure of a contact on the touch sensing surface. The intensity of a contact has a range of values that includes at least four distinct values, and more typically, hundreds (e.g., at least 256) of distinct values. The intensity of a contact is optionally determined (or measured) using a variety of techniques and a variety of sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch sensing surface are optionally used to measure the force at various points on the touch sensing surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine the estimated force of the contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch sensing surface. Alternatively, the size and / or change thereof of the contact area detected on the touch sensing surface, the capacitance and / or change thereof of the touch sensing surface in proximity to the contact, and / or the resistance and / or change thereof of the touch sensing surface in proximity to the contact are optionally used as surrogates for the force or pressure of a contact on the touch sensing surface. In some implementations, alternative measurements of the force or pressure of a contact are used directly to determine whether the alternative measurement exceeds an intensity threshold (e.g., the intensity threshold is described in units corresponding to the alternative measurement). In some implementations, a proxy measurement of the contact force or pressure is converted to an estimated value of the force or pressure, and the estimated value of the force or pressure is used to determine whether the estimated value exceeds an intensity threshold (e.g., the intensity threshold is a pressure threshold measured in units of pressure). By using the intensity of a contact as an attribute of user input, a user can access additional device functions that may not otherwise be accessible to the user on a reduced-size device with a limited implementation area for displaying affordances (e.g., on a touch sensing display), and / or receive user input (e.g., via a touch sensing display, a touch sensing surface, or a physical / mechanical control such as a knob or button).
[0055] As used in this specification and the claims, the term "haptic output" refers to the physical displacement of the device relative to its previous position, the physical displacement of a component of the device (e.g., a touch-sensing surface) relative to another component of the device (e.g., the housing), or the displacement of a component relative to the center of mass of the device, which is to be detected by the user's sense of touch. For example, in a situation where the device or a component of the device is in contact with a touch-sensitive surface of the user (e.g., the finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or the component of the device. For example, the movement of a touch-sensing surface (e.g., a touch-sensing display or a trackpad) can optionally be interpreted by the user as a "down click" or "up click" of a physical actuator button. In some cases, the user may feel a tactile sensation such as a "down click" or "up click" even when there is no movement of the physical actuator button associated with the touch-sensing surface physically pressed (e.g., displaced) by the user's action. As another example, the movement of a touch-sensing surface can optionally be interpreted or perceived by the user as the "roughness" of the touch-sensing surface even when there is no change in the smoothness of the touch-sensing surface. Such interpretation of touch by the user depends on the user's individual sensory perception, but there are many sensory perceptions of touch that are common to a majority of users. Therefore, when a haptic output is described as corresponding to a particular sensory perception of the user (e.g., "up click", "down click", "roughness"), unless otherwise stated, the generated haptic output corresponds to the physical displacement of the device or a component of the device that produces the described sensory perception of a typical (or average) user.
[0056] Device 200 is merely an example of a portable multifunctional device. Device 200 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of those components. It should be understood that the various components shown in FIG. 2A are implemented as hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.
[0057] Memory 202 includes one or more computer-readable storage media. This computer-readable storage media is, for example, tangible and non-transitory. Memory 202 includes high-speed random access memory and also includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid state memory devices. Memory controller 222 controls access to memory 202 by other components of device 200.
[0058] In some embodiments, the non-transitory computer-readable storage media of memory 202 is used to store instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other system capable of fetching and executing instructions from that instruction execution system, apparatus, or device (e.g., for executing aspects of the processes described below). In other embodiments, the instructions (e.g., for executing aspects of the processes described below) are stored in a non-transitory computer-readable storage media of server system 108 (not shown) or are divided between the non-transitory computer-readable storage media of memory 202 and the non-transitory computer-readable storage media of server system 108.
[0059] The peripheral device interface 218 is used to couple the input and output peripheral devices of this device to the CPU 220 and the memory 202. One or more processors 220 operate or execute various software programs and / or instruction sets stored in the memory 202 to perform various functions for the device 200 and process data. In some embodiments, the peripheral device interface 218, the CPU 220, and the memory controller 222 are implemented on a single chip such as the chip 204. In some other embodiments, they are implemented on separate chips.
[0060] The RF (radio frequency) circuit 208 transmits and receives RF signals, also called electromagnetic signals. The RF circuit 208 converts electrical signals into electromagnetic signals or vice versa and communicates with a communication network and other communication devices via electromagnetic signals. The RF circuit 208 optionally includes well-known circuits for performing these functions, such as, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, and the like. The RF circuit 208 optionally communicates wirelessly with a network such as the Internet, also called the World Wide Web (WWW), an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN), as well as with other devices. The RF circuit 208 optionally includes well-known circuits for detecting a near field communication (NFC) field by, for example, a short-range communication radio. Wireless communication optionally includes, but is not limited to only, Global System for Mobile Communications (GSM) for mobile communication, Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPDA), Long Termevolution, LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE (registered trademark)), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX (registered trademark), protocols for email (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including communication protocols not yet developed as of the filing date of this specification, using any of a plurality of communication standards, protocols, and technologies.
[0061] The audio circuit 210, the speaker 211, and the microphone 213 provide an audio interface between the user and the device 200. The audio circuit 210 receives audio data from the peripheral device interface 218, converts this audio data into an electrical signal, and transmits this electrical signal to the speaker 211. The speaker 211 converts the electrical signal into human audible sound waves. Also, the audio circuit 210 receives the electrical signal converted from sound waves by the microphone 213. The audio circuit 210 converts the electrical signal into audio data and transmits this audio data to the peripheral device interface 218 for processing. The audio data is obtained by the peripheral device interface 218 from the memory 202 and / or the RF circuit 208, and / or transmitted to the memory 202 and / or the RF circuit 208. In some embodiments, the audio circuit 210 also includes a headset jack (e.g., 312 of FIG. 3). The headset jack provides an interface between the audio circuit 210 and a detachable audio input / output peripheral device such as an output-only headset or a headset with both output (e.g., mono or stereo headphones) and input (e.g., microphone).
[0062] The I / O subsystem 206 couples input / output peripheral devices on the device 200, such as the touch screen 212 and other input control devices 216, to the peripheral device interface 218. The I / O subsystem 206 optionally includes a display controller 256, an optical sensor controller 258, an intensity sensor controller 259, a haptic feedback controller 261, and one or more input controllers 260 for other input or control devices. The one or more input controllers 260 receive electrical signals from and transmit electrical signals to other input control devices 216. The other input control devices 216 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, the input controller(s) 260 are optionally coupled to (or not coupled to any of) a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. One or more buttons (e.g., 308 in FIG. 3) optionally include up / down buttons for volume control of the speaker 211 and / or the microphone 213. One or more buttons optionally include a push button (e.g., 306 in FIG. 3).
[0063] Quickly pressing a push button unlocks the touch screen 212 or initiates a process of using gestures on the touch screen to unlock the device, as described in U.S. Patent Application No. 11 / 322,549, filed on December 23, 2005, and U.S. Patent No. 7,657,849, "Unlocking a Device by Performing Gestures on an Unlock Image," the entire contents of which are incorporated herein by reference. Pressing a push button (e.g., 306) for an extended period of time turns the device 200 on or off. The user can customize the functions of the one or more buttons. The touch screen 212 is used to implement virtual or soft buttons and one or more soft keyboards.
[0064] The touch-sensitive display 212 provides an input interface and an output interface between the device and the user. The display controller 256 receives electrical signals from the touch screen 212 and / or transmits electrical signals to the touch screen 212. The touch screen 212 displays visual output to the user. The visual output includes graphics, text, icons, videos, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output corresponds to user interface objects.
[0065] The touch screen 212 has a touch-sensitive surface, sensor, or set of sensors that accepts input from the user based on tactile and / or haptic contact. The touch screen 212 and the display controller 256 (along with any associated modules and / or instruction sets in the memory 202) detect contact (and any movement or interruption of the contact) on the touch screen 212 and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 212. In an exemplary embodiment, the point of contact between the touch screen 212 and the user corresponds to the user's finger.
[0066] The touch screen 212 uses LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, although in other embodiments other display technologies may be used. The touch screen 212 and the display controller 256 detect contact and its movement or break using any of a plurality of touch sensing technologies currently known or to be developed in the future, including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more contact points using the touch screen 212. In an exemplary embodiment, a projected mutual capacitance sensing technology such as that found in the iPhone (registered trademark) and iPod Touch (registered trademark) from Apple Inc. of Cupertino, California, is used.
[0067] The touch sensing display of some embodiments of the touch screen 212 is similar to the multi-touch sensing touch pads described in U.S. Patent Nos. 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman) and / or U.S. Patent Publication 2002 / 0015024 (A1), each of which is incorporated herein by reference in its entirety. However, the touch screen 212 displays visual output from the device 200, whereas the touch sensing touch pad does not provide visual output.
[0068] The touch sensing display in some embodiments of the touch screen 212 is as described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, filed May 2, 2006, "Multipoint Touch Surface Controller"; (2) U.S. Patent Application No. 10 / 840,862, filed May 6, 2004, "Multipoint Touchscreen"; (3) U.S. Patent Application No. 10 / 903,964, filed Jul. 30, 2004, "Gestures For Touch Sensitive Input Devices"; (4) U.S. Patent Application No. 11 / 048,264, filed Jan. 31, 2005, "Gestures For Touch Sensitive Input Devices"; (5) U.S. Patent Application No. 11 / 038,590, filed Jan. 18, 2005, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices"; (6) U.S. Patent Application No. 11 / 228,758, filed Sep. 16, 2005, "Virtual Input Device Placement On A Touch Screen User Interface"; (7) U.S. Patent Application No. 11 / 228,700, filed Sep. 16, 2005, "Operation Of A Computer With A Touch Screen Interface"; (8) U.S. Patent Application No. 11 / 228,737, filed Sep. 16, 2005, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard"; and (9) U.S. Patent Application No. 11 / 367,749, filed Mar. 3, 2006, "Multi-Functional Hand-Held Device". All of these applications are hereby incorporated by reference in their entirety.
[0069] The touch screen 212 has, for example, a video resolution exceeding 100 dpi. In some embodiments, the touch screen has a video resolution of about 160 dpi. The user touches the touch screen 212 using a suitable object or appendage such as a stylus, finger, etc. In some embodiments, the user interface is designed to operate primarily using finger-based contact and gestures, although this may be less accurate than stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts rough input by a finger into an accurate pointer / cursor position or command for performing the action desired by the user.
[0070] In some embodiments, in addition to the touch screen, the device 200 includes a touch pad (not shown) that activates or deactivates certain functions. In some embodiments, the touch pad, unlike the touch screen, is a touch-sensitive area of the device that does not display a visual output. The touch pad is a separate touch-sensitive surface from the touch screen 212 or an extension of the touch-sensitive surface formed by the touch screen.
[0071] The device 200 also includes a power system 262 that supplies power to various components. The power system 262 includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a charging system, a power outage detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0072] Device 200 also includes one or more optical sensors 264. FIG. 2A shows an optical sensor coupled to an optical sensor controller 258 within I / O subsystem 206. Optical sensor 264 includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Optical sensor 264 receives light from the environment projected through one or more lenses and converts that light into data representative of an image. In conjunction with imaging module 243 (also referred to as a camera module), optical sensor 264 captures a still image or video. In some embodiments, the optical sensor is disposed on the back of device 200 opposite touch screen display 212 on the front of the device such that the touch screen display is used as a viewfinder for still image and / or video capture. In some embodiments, the optical sensor is disposed on the front of the device such that a user's image for a video conference is captured while the user views other video conference participants on the touch screen display. In some embodiments, the position of optical sensor 264 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), and thus a single optical sensor 264 is used with the touch screen display for both video conferencing and still image and / or video capture.
[0073] Device 200 also optionally includes one or more contact intensity sensors 265. FIG. 2A shows a contact intensity sensor coupled to an intensity sensor controller 259 within the I / O subsystem 206. The contact intensity sensor 265 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric force sensors, optical force sensors, capacitive touch sensing surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of contact on a touch sensing surface). The contact intensity sensor 265 receives contact intensity information (e.g., pressure information, or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with, or proximate to, a touch sensing surface (e.g., touch sensing display system 212). In some embodiments, at least one contact intensity sensor is disposed on the back of the device 200, opposite a touch screen display 212 disposed on the front of the device 200.
[0074] Device 200 also includes one or more proximity sensors 266. FIG. 2A shows a proximity sensor 266 coupled to the peripheral device interface 218. Alternatively, the proximity sensor 266 is coupled to an input controller 260 within the I / O subsystem 206. The proximity sensor 266 functions as described in U.S. Patent Application Nos. 11 / 241,839, "Proximity Detector In Handheld Device"; 11 / 240,788, "Proximity Detector In Handheld Device"; 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output"; 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices"; and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals". In some embodiments, when a multifunctional device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor turns off and disables the touch screen 212.
[0075] Device 200 also optionally includes one or more haptic output generators 267. FIG. 2A shows a haptic output generator coupled to a haptic feedback controller 261 within I / O subsystem 206. The haptic output generator 267 optionally includes one or more electroacoustic devices, such as speakers or other audio components, and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components that convert an electrical signal into a haptic output) on the device. The contact intensity sensor 265 receives haptic feedback generation instructions from the haptic feedback module 233 and generates a haptic output on device 200 that can be sensed by a user of device 200. In some embodiments, at least one haptic output generator is juxtaposed with, or proximate to, a touch sensing surface (e.g., touch sensing display system 212) and optionally generates a haptic output by moving the touch sensing surface in a vertical direction (e.g., in / out of the surface of device 200) or in a horizontal direction (e.g., back and forth within the same plane as the surface of device 200). In some embodiments, at least one haptic output generator sensor is disposed on the back of device 200, which is opposite the touch screen display 212 disposed on the front of device 200.
[0076] Device 200 also includes one or more accelerometers 268. FIG. 2A shows an accelerometer 268 coupled to the peripheral device interface 218. Alternatively, the accelerometer 268 is coupled to an input controller 260 within the I / O subsystem 206. The accelerometer 268 operates, for example, as described in "Acceleration-based Theft Detection for Portable Electronic Devices" of U.S. Patent Publication No. 20050190059 and "Acceleration-based Theft Detection System for Portable Electronic Devices" of U.S. Patent Publication No. 20060017692, both of which are hereby incorporated by reference in their entirety. In some embodiments, information is displayed on the touch screen display in a portrait or landscape orientation based on an analysis of data received from one or more accelerometers. Device 200 optionally includes, in addition to the accelerometer(s) 268, a magnetometer (not shown), and a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information regarding the location and orientation (e.g., portrait or landscape) of device 200.
[0077] In some embodiments, the software components stored in memory 202 include an operating system 226, a communication module (or instruction set) 228, a touch / motion module (or instruction set) 230, a graphics module (or instruction set) 232, a text input module (or instruction set) 234, a Global Positioning System (GPS) module (or instruction set) 235, a digital assistant client module 229, and an application (or instruction set) 236. Further, memory 202 stores data and models, such as user data and model 231. Further, in some embodiments, as shown in FIGS. 2A and 4, memory 202 (FIG. 2A) or memory 470 (FIG. 4) stores a device / global internal state 257. The device / global internal state 257 includes an active application state indicating which application is active if there is a currently active application, a display state indicating which application, view, or other information occupies various regions of the touch screen display 212, a sensor state including information obtained from various sensors and input control devices 216 of the device, and one or more of location information regarding the location and / or orientation of the device.
[0078] The operating system 226 (e.g., an embedded operating system such as Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitate communication between various hardware components and software components.
[0079] The communication module 228 facilitates communication with other devices via one or more external ports 224 and also includes various software components for processing data received by the RF circuit 208 and / or the external ports 224. The external ports 224 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) are adapted to couple to other devices either directly or indirectly via a network (e.g., the Internet, a wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as or similar to and / or compatible with the 30-pin connector used on iPod (registered trademark) devices (a trademark of Apple Inc.).
[0080] The contact / motion module 230 optionally detects contact with the touch screen 212 and other touch-sensitive devices (e.g., a touch pad or a physical click wheel) (in cooperation with the display controller 256). The contact / motion module 230 performs various operations related to the detection of contact, such as determining whether contact has occurred (e.g., detecting a finger-down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or an alternative to the force or pressure of the contact), determining whether there is movement of the contact, tracking movement across the touch-sensitive surface (e.g., detecting one or more events of dragging a finger), and determining whether the contact has ceased (e.g., detecting a finger-up event or an interruption of the contact), by including various software components. The contact / motion module 230 receives contact data from the touch-sensitive surface. Determining the movement of the contact point, represented by a series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. These operations are optionally applicable to a single contact (e.g., contact by one finger) or multiple simultaneous contacts (e.g., "multi-touch" / contact by multiple fingers). In some embodiments, the contact / motion module 230 and the display controller 256 detect contact on a touch pad.
[0081] In some embodiments, the contact / motion module 230 uses a set of one or more intensity thresholds to determine whether an action has been performed by the user (e.g., to determine whether the user has "clicked" on an icon). In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (e.g., the intensity thresholds can be adjusted without changing the physical hardware of the device 200, rather than being determined by the activation thresholds of specific physical actuators). For example, the mouse "click" threshold of a trackpad or touch screen display can be set to any of a wide range of defined thresholds without changing the trackpad or touch screen display hardware. Additionally, in some implementations, the user of the device is provided with software settings to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or adjusting multiple intensity thresholds at once by system-level click "intensity" parameters).
[0082] The contact / motion module 230 optionally detects gesture inputs from the user. Different gestures on the touch-sensing surface have different contact patterns (e.g., the detected movement, timing, and / or intensity of the contact is different). Thus, gestures are optionally detected by detecting a specific contact pattern. For example, detecting a finger tap gesture includes detecting a finger down event followed by a finger up (lift off) event at the same position (or substantially the same position) as the finger down event (e.g., the position of an icon). As another example, detecting a finger swipe gesture on the touch-sensing surface includes detecting a finger down event followed by one or more finger drag events followed by a finger up (lift off) event.
[0083] The graphic module 232 includes various known software components for rendering and displaying graphics on the touch screen 212 or other display, including components for changing the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). As used herein, the term "graphic" includes any object that can be displayed to the user, including, but not limited to, text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0084] In some embodiments, the graphic module 232 stores data representing the graphics that will be used. Each graphic is optionally assigned a corresponding code. The graphic module 232 receives one or more codes specifying the graphics to be displayed, along with coordinate data and other graphic property data as needed, from an application or the like, and then generates the screen image data to be output to the display controller 256.
[0085] The haptic feedback module 233 includes various software components for generating the instructions used by the haptic output generator(s) 267 to generate haptic output at one or more locations on the device 200 in response to user interaction with the device 200.
[0086] The text input module 234, which is a component of the graphic module 232, provides a soft keyboard for entering text in some embodiments in various applications (e.g., contacts 237, email 240, IM 241, browser 247, and any other application that requires text input).
[0087] The GPS module 235 determines the location of the device and provides this information for use within various applications (e.g., to the telephone 238 for use in location-based dialing, to the camera 243 as picture / video metadata, and to applications that provide location-based services such as a weather widget, a local yellow pages widget, and a map / navigation widget).
[0088] The digital assistant client module 229 includes various client-side digital assistant instructions for providing client-side functionality of the digital assistant. For example, the digital assistant client module 229 can accept voice input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces of the portable multifunction device 200 (e.g., the microphone 213, the accelerometer(s) 268, the touch-sensing display system 212, the optical sensor(s) 264, other input control devices 216, etc.). The digital assistant client module 229 can also provide output such as audio (e.g., speech output), visual, and / or haptic output through various output interfaces of the portable multifunction device 200 (e.g., the speaker 211, the touch-sensing display system 212, the haptic output generator(s) 267, etc.). For example, the output is provided as voice, sound, alerts, text messages, menus, graphics, video, animations, vibrations, and / or combinations of two or more of the above. During operation, the digital assistant client module 229 communicates with the DA server 106 using the RF circuit 208.
[0089] The user data and model 231 include various data associated with the user (e.g., user-specific vocabulary data, user preference data, pronunciation of user-specified names, data from the user's electronic address book, to-do lists, shopping lists, etc.) for providing client-side functions of the digital assistant. Furthermore, the user data and model 231 include various models (e.g., speech recognition models, statistical language models, natural language processing models, ontologies, task flow models, service models, etc.) for processing user input and determining user intent.
[0090] In some embodiments, the digital assistant client module 229 utilizes various sensors, subsystems, and peripheral devices of the portable multifunctional device 200 to collect additional information from the surrounding environment of the portable multifunctional device 200, thereby establishing context associated with the user, the current user interaction, and / or the current user input. In some embodiments, the digital assistant client module 229 provides context information or a subset thereof to the DA server 106 along with the user input to assist in inferring the user's intent. In some embodiments, the digital assistant also uses the context information to determine how to prepare and deliver an output to the user. The context information is referred to as context data.
[0091] In some embodiments, context information involving user input includes sensor information, such as lighting, ambient noise, ambient temperature, images or videos of the ambient environment, etc. In some embodiments, context information may also include the physical state of the device, such as the orientation of the device, the location of the device, the temperature of the device, the power level, the speed, the acceleration, the movement pattern, the cellular signal strength, etc. In some embodiments, information regarding the software state of DA server 106, such as running processes, installed programs, past and current network activities, background services, error logs, resource usage, etc., and information regarding the software state of portable multifunctional device 200 are provided to DA server 106 as context information associated with the user input.
[0092] In some embodiments, digital assistant client module 229 selectively provides information (such as user data 231) stored on portable multifunctional device 200 in response to requests from DA server 106. In some embodiments, digital assistant client module 229 also draws additional input from the user via a natural language dialog or other user interface in response to requests by DA server 106. Digital assistant client module 229 passes that additional input to DA server 106 to assist DA server 106 in inferring and / or fulfilling the user's intent expressed within the user request.
[0093] A more detailed description of the digital assistant will be provided below with reference to FIGS. 7A - 7C. It should be recognized that digital assistant client module 229 may include any number of sub - modules of digital assistant module 726 described below.
[0094] Application 236 includes the following modules (or sets of instructions), or subsets or supersets thereof. ● Contact module 237 (which may also be referred to as an address book or contact list), ● Telephone module 238, ● Video conferencing module 239, ● Email client module 240, ● Instant messaging (IM) module 241, ● Training support module 242, ● Camera module 243 for still images and / or videos, ● Image management module 244, ● Video player module, ● Music player module, ● Browser module 247, ● Calendar module 248, ● Widget module 249, which in some embodiments includes one or more of weather widget 249-1, stock price widget 249-2, calculator widget 249-3, alarm clock widget 249-4, dictionary widget 249-5, and other widgets obtained by the user as well as user-created widget 249-6, ● Widget creator module 250 for creating user-created widget 249-6, ● Search module 251, ● Integrated video and music player module 252 that integrates the video player module and the music player module, ● Note module 253, ● Map module 254, and / or ● Online video module 255.
[0095] Examples of other applications 236 stored in memory 202 include other word processing applications, other image editing applications, drawing applications, presentation applications, Java-compatible applications, encryption, digital rights management, speech recognition, and speech replication.
[0096] In conjunction with the touch screen 212, display controller 256, contact / motion module 230, graphic module 232, and text input module 234, the contact module 237 is used to manage an address book or contact list (e.g., stored in the application internal state 292 of the contact module 237 in memory 202 or memory 470), including adding a name (singular or plural) to the address book, deleting a name (singular or plural) from the address book, associating a phone number (singular or plural), email address (singular or plural), physical address (singular or plural), or other information with a name. This includes associating an image with a name, classifying and sorting names, and providing a phone number or email address to initiate and / or facilitate communication by phone 238, video conferencing module 239, email 240, or IM 241.
[0097] In conjunction with the RF circuit 208, audio circuit 210, speaker 211, microphone 213, touch screen 212, display controller 256, contact / motion module 230, graphic module 232, and text input module 234, the phone module 238 is used to input a string corresponding to a phone number, access one or more phone numbers in the contact module 237, modify the input phone number, dial an individual phone number, conduct a conversation, and disconnect or hang up the phone when the conversation is complete. Thus, wireless communication uses any of a plurality of communication standards, protocols, and technologies.
[0098] The video conferencing module 239 includes executable instructions to initiate, execute, and end a video conference between the user and one or more other participants according to the user's instructions, in conjunction with the RF circuit 208, audio circuit 210, speaker 211, microphone 213, touch screen 212, display controller 256, optical sensor 264, optical sensor controller 258, contact / motion module 230, graphic module 232, text input module 234, contact module 237, and phone module 238.
[0099] The email client module 240 includes executable instructions for creating, sending, receiving, and managing emails in response to user commands in cooperation with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphic module 232, and text input module 234. In cooperation with the image management module 244, the email client module 240 makes it very easy to create and send emails with still or moving images captured by the camera module 243.
[0100] The instant messaging module 241 includes executable instructions for inputting a series of characters corresponding to an instant message, correcting previously input characters, (e.g., using the Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for phone communication-based instant messages, or XMPP, SIMPLE, or IMPS for Internet-based instant messages) sending individual instant messages, receiving instant messages, and viewing received instant messages in cooperation with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphic module 232, and text input module 234. In some embodiments, the sent and / or received instant messages include graphics, photos, audio files, video files, and / or other attachments supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both phone communication-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0101] In cooperation with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphic module 232, text input module 234, GPS module 235, map module 254, and music player module, the training support module 242 includes executable instructions that create a training (e.g., having time, distance, and / or calorie burn goals), communicate with a training sensor (sports device), receive training sensor data, calibrate sensors used to monitor the training, select and play music for the training, and display, store, and transmit training data.
[0102] In cooperation with the touch screen 212, display controller 256, optical sensor(s) 264, optical sensor controller 258, contact / motion module 230, graphic module 232, and image management module 244, the camera module 243 includes executable instructions for capturing a still image or video (including a video stream) and storing it in the memory 202, modifying the characteristics of a still image or video, or deleting a still image or video from the memory 202.
[0103] In cooperation with the touch screen 212, display controller 256, contact / motion module 230, graphic module 232, text input module 234, and camera module 243, the image management module 244 includes executable instructions for arranging, modifying (e.g., editing), or otherwise operating on, labeling, deleting, presenting (e.g., in a digital slide show or album), and storing still images and / or videos.
[0104] The browser module 247 includes executable instructions for browsing the Internet according to user instructions, including searching for, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to the web page, in cooperation with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234.
[0105] The calendar module 248 includes executable instructions for creating, displaying, modifying, and storing a calendar and data associated with the calendar (e.g., calendar items, to-do lists, etc.) according to user instructions, in cooperation with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, email client module 240, and browser module 247.
[0106] In conjunction with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphic module 232, text input module 234, and browser module 247, the widget module 249 is a mini-application that can be downloaded and used by a user (e.g., weather widget 249-1, stock price widget 249-2, calculator widget 249-3, alarm clock widget 249-4, and dictionary widget 249-5), or can be created by the user (e.g., user-created widget 249-6). In some embodiments, the widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, the widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! widget).
[0107] In conjunction with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphic module 232, text input module 234, and browser module 247, the widget creation module 250 is used by a user to create a widget (e.g., change a user-specified portion of a web page into a widget).
[0108] The search module 251 includes executable instructions for searching for characters, music, sound, images, videos, and / or other files in the memory 202 that match one or more search criteria (e.g., one or more user-specified search terms) according to a user's command in cooperation with the touch screen 212, display controller 256, contact / motion module 230, graphic module 232, and text input module 234.
[0109] The video and music player module 252, in conjunction with the touch screen 212, display controller 256, contact / motion module 230, graphic module 232, audio circuit 210, speaker 211, RF circuit 208, and browser module 247, includes executable instructions that enable a user to download and play recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, and executable instructions for displaying, presenting, or otherwise playing videos (e.g., on the touch screen 212 or on an external display connected via the external port 224). In some embodiments, the device 200 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).
[0110] The memo module 253, in conjunction with the touch screen 212, display controller 256, contact / motion module 230, graphic module 232, and text input module 234, includes executable instructions for creating and managing memos, to-do lists, etc. according to the user's instructions.
[0111] In conjunction with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphic module 232, text input module 234, GPS module 235, and browser module 247, the map module 254 is used to receive, display, modify, and store maps and data associated with the maps (e.g., driving directions, data about stores, specific locations or other locations near a particular place, and other location-based data) according to the user's instructions.
[0112] The online video module 255, in cooperation with the touch screen 212, the display controller 256, the touch / motion module 230, the graphics module 232, the audio circuit 210, the speaker 211, the RF circuit 208, the text input module 234, the email client module 240, and the browser module 247, includes instructions that enable a user to access a particular online video, browse a particular online video, receive it (e.g., by streaming and / or downloading), play it (e.g., on the touch screen or on an external display connected via the external port 224), send an email having a link to a particular online video, and perform other management of online videos in one or more file formats such as H.264. In some embodiments, instead of the email client module 240, the instant messaging module 241 is used to send a link to a particular online video. Additional explanation of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed Jun. 20, 2007, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos", and U.S. Patent Application No. 11 / 968,067, filed Dec. 31, 2007, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos", the entire contents of which are incorporated herein by reference.
[0113] The modules and applications identified above each correspond to a set of executable instructions that perform one or more of the functions described above and the methods described in this application (e.g., the computer-executed methods and other information processing methods described herein). Since these modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules, various subsets of these modules can be combined or otherwise reconfigured in various embodiments. For example, a video player module can be combined with a music player module into a single module (e.g., video and music player module 252, FIG. 2A). In some embodiments, memory 202 stores a subset of the modules and data structures identified above. Further, memory 202 stores additional modules and data structures not described above.
[0114] In some embodiments, device 200 is a device in which the operation of a set of default functions in the device is performed only via a touch screen and / or a touch pad. By using a touch screen and / or a touch pad as the primary input control device for the operation of device 200, the number of physical input control devices (push buttons, dials, etc.) on device 200 is reduced.
[0115] The set of default functions that are performed only through the touch screen and / or the touch pad optionally includes navigation between user interfaces. In some embodiments, the touch pad navigates device 200 from any user interface displayed on device 200 to the main menu, home menu, or root menu when touched by the user. In such embodiments, the "menu button" is implemented using the touch pad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touch pad.
[0116] FIG. 2B is a block diagram showing exemplary components for event processing, according to some embodiments. In some embodiments, memory 202 (FIG. 2A) or memory 470 (FIG. 4) includes an event sorter 270 (e.g., within operating system 226) and individual applications 236-1 (e.g., any of the aforementioned applications 237-251, 255, 480-490).
[0117] The event sorter 270 receives event information and determines an application 236-1 to which the event information is to be delivered, and an application view 291 of the application 236-1. The event sorter 270 includes an event monitor 271 and an event dispatcher module 274. In some embodiments, the application 236-1 includes an application internal state 292 that indicates the current application view(s) displayed on the touch-sensitive display 212 when the application is active or running. In some embodiments, the device / global internal state 257 is used by the event sorter 270 to determine which application(s) is / are currently active, and the application internal state 292 is used by the event sorter 270 to determine the application view 291 to which the event information is to be delivered.
[0118] In some embodiments, the application internal state 292 includes additional information such as resume information to be used when the application 236-1 resumes execution, user interface state information indicating or ready to display information being displayed by the application 236-1, a state queue that enables the user to return to a previous state or view of the application 236-1, and a redo / undo queue of previous actions performed by the user.
[0119] The event monitor 271 receives event information from the peripheral device interface 218. The event information includes information regarding sub-events (e.g., a user touch as part of a multi-touch gesture on the touch-sensitive display 212). The peripheral device interface 218 transmits information received from the I / O subsystem 206, or sensors such as the proximity sensor 266, the accelerometer(s) 268, and / or the microphone 213 (via the audio circuitry 210). The information that the peripheral device interface 218 receives from the I / O subsystem 206 includes information from the touch-sensitive display 212 or a touch-sensitive surface.
[0120] In some embodiments, the event monitor 271 transmits requests to the peripheral device interface 218 at predetermined intervals. In response, the peripheral device interface 218 transmits event information. In other embodiments, the peripheral device interface 218 transmits event information only when there is an important event (e.g., receipt of an input that exceeds a predetermined noise threshold and / or exceeds a predetermined duration).
[0121] In some embodiments, the event sorter 270 also includes a hit view determination module 272 and / or an active event recognition unit determination module 273.
[0122] The hit view determination module 272 provides software procedures for determining where in one or more views a sub-event occurs when the touch-sensitive display 212 is displaying two or more views. A view is composed of control devices and other elements that a user can view on the display.
[0123] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as an application view or a user interface window, in which information is displayed and touch-based gestures occur. The application view (of an individual application) in which a touch is detected corresponds to a program level within the program of the application or the view hierarchy. For example, the lowest level view in which a touch is detected is called the hit view, and the set of events recognized as valid input is determined based at least in part on the hit view of the initial touch that starts a touch-based gesture.
[0124] The hit view determination module 272 receives information related to sub-events of a touch-based gesture. When the application has a plurality of hierarchically structured views, the hit view determination module 272 identifies the hit view as the lowest level view within the hierarchy in which the sub-event is to be processed. In most situations, the hit view is the lowest level view at which the start sub-event (e.g., the first sub-event in a series of sub-events that form an event or potential event) occurs. Once the hit view is identified by the hit view determination module 272, the hit view typically receives all sub-events related to the same touch or input source that was identified as the hit view.
[0125] The active event recognition unit determination module 273 determines which view(s) within the view hierarchy should receive a particular series of sub-events. In some embodiments, the active event recognition unit determination module 273 determines that only the hit view should receive a particular series of sub-events. In other embodiments, the active event recognition unit determination module 273 determines that all views including the physical location of the sub-events are views actively involved, and thus determines that all views actively involved should receive a particular series of sub-events. In other embodiments, even if a touch sub-event is completely limited to an area associated with one particular view, the upper-level views within the hierarchy remain views that are actively involved.
[0126] The event dispatcher module 274 dispatches event information to the event recognition unit (e.g., event recognition unit 280). In embodiments including the active event recognition unit determination module 273, the event dispatcher module 274 distributes event information to the event recognition unit determined by the active event recognition unit determination module 273. In some embodiments, the event dispatcher module 274 stores the event information retrieved by the individual event receiver 282 in the event queue.
[0127] In some embodiments, the operating system 226 includes the event sorter 270. Alternatively, the application 236-1 includes the event sorter 270. In still other embodiments, the event sorter 270 is a stand-alone module or part of another module stored in the memory 202 such as the touch / motion module 230.
[0128] In some embodiments, application 236-1 includes a plurality of event processing units 290 and one or more application views 291, each including instructions to process touch events occurring within an individual view of the application's user interface. Each application view 291 of application 236-1 includes one or more event recognition units 280. Typically, an individual application view 291 includes a plurality of event recognition units 280. In other embodiments, one or more of the event recognition units 280 are part of a separate module, such as a user interface kit (not shown) or a higher-level object from which application 236-1 inherits methods and other characteristics. In some embodiments, an individual event processing unit 290 includes one or more of event data 279 received from data update unit 276, object update unit 277, GUI update unit 278, and / or event sorting unit 270. The event processing unit 290 utilizes or calls data update unit 276, object update unit 277, or GUI update unit 278 to update the internal state 292 of the application. Alternatively, one or more of the application views 291 include one or more respective event processing units 290. Also, in some embodiments, one or more of data update unit 276, object update unit 277, and GUI update unit 278 are included within an individual application view 291.
[0129] An individual event recognition unit 280 receives event information (e.g., event data 279) from event sorting unit 270 and identifies events from the event information. The event recognition unit 280 includes an event receiving unit 282 and an event comparing unit 284. In some embodiments, the event recognition unit 280 includes at least a subset of metadata 283 and event delivery instructions 288 (including sub-event delivery instructions).
[0130] The event receiving unit 282 receives event information from the event sorting unit 270. The event information includes sub-events, for example, information about a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information such as the location of the sub-event. When the sub-event is related to the movement of a touch, the event information also includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the posture of the device).
[0131] The event comparison unit 284 compares the event information with the definition of a predefined event or sub-event, and based on the comparison, determines an event or sub-event, or determines or updates the state of an event or sub-event. In some embodiments, the event comparison unit 284 includes an event definition 286. The event definition 286 includes definitions of events (e.g., a predefined series of sub-events) such as event 1 (287-1) and event 2 (287-2). In some embodiments, the sub-events within an event (e.g., event 1 (287-1) and event 2 (287-2)) include, for example, a touch start, a touch end, a touch movement, a touch cancellation, and multiple touches. In one example, the definition of event 1 (287-1) is a double tap on a displayed object. The double tap includes, for example, a first touch (touch start) on the displayed object for a predetermined stage, a first lift-off (touch end) for the predetermined stage, a second touch (touch start) on the displayed object for the predetermined stage, and a second lift-off (touch end) for the predetermined stage. In another example, the definition of event 2 (287-2) is a drag on a displayed object. The drag includes, for example, a touch (or contact) on the displayed object for a predetermined stage, a movement of the touch across the touch-sensitive display 212, and a lift-off of the touch (touch end). In some embodiments, the event also includes information about one or more associated event processing units 290.
[0132] In some embodiments, the event definition 286 includes the definition of events for individual user interface objects. In some embodiments, the event comparison unit 284 performs a hit test to determine which user interface object is associated with the sub - event. For example, within an application view where three user interface objects are displayed on the touch - sensitive display 212, when a touch is detected on the touch - sensitive display 212, the event comparison unit 284 performs a hit test to determine which of the three user interface objects is associated with the touch (sub - event). If each of the displayed objects is associated with an individual event processing unit 290, the event comparison unit determines which event processing unit 290 should be activated using the result of the hit test. For example, the event comparison unit 284 selects the event processing unit associated with the sub - event and object that triggered the hit test.
[0133] In some embodiments, the definition of an individual event (e.g., event 1 (287 - 1) or event 2 (287 - 2)) also includes a delay action that delays the delivery of event information until it is determined whether a series of sub - events corresponds to the event type of the event recognition unit.
[0134] If the individual event recognition unit 280 determines that a series of sub - events does not match any of the events in the event definition 286, the individual event recognition unit 280 enters a state of event impossible, event failure, or event end, and then ignores the next sub - event of the touch - based gesture. In this situation, if there are other event recognition units that remain active for the hit view, those event recognition units continue to track and process the sub - events of the ongoing touch - based gesture.
[0135] In some embodiments, the individual event recognition unit 280 includes metadata 283 having configurable properties, flags, and / or lists indicating how the event distribution system should actively participate in the event recognition unit that performs sub-event distribution. In some embodiments, the metadata 283 includes configurable properties, flags, and / or lists indicating how the event recognition units interact with each other or how they can interact with each other. In some embodiments, the metadata 283 includes configurable properties, flags, and / or lists indicating whether sub-events are distributed at various levels in the view hierarchy or program hierarchy.
[0136] In some embodiments, the individual event recognition unit 280 activates the event processing unit 290 associated with the event when one or more specific sub-events of the event are recognized. In some embodiments, the individual event recognition unit 280 distributes the event information associated with the event to the event processing unit 290. Activating the event processing unit 290 is separate from sending (and deferring sending) sub-events to individual hit views. In some embodiments, the event recognition unit 280 sets a flag associated with the recognized event, and the event processing unit 290 associated with the flag catches the flag and executes a predefined process.
[0137] In some embodiments, the event distribution command 288 includes a sub-event distribution command that distributes event information about sub-events without activating the event processing unit. Instead, the sub-event distribution command distributes the event information to the event processing unit associated with a series of sub-events or to the view actively involved. The event processing unit associated with a series of sub-events or the view actively involved receives the event information and executes a predetermined process.
[0138] In some embodiments, the data update unit 276 creates and updates data used in the application 236-1. For example, the data update unit 276 updates the phone numbers used in the contact module 237 or stores the video files used in the video player module. In some embodiments, the object update unit 277 creates and updates objects used in the application 236-1. For example, the object update unit 277 creates a new user interface object or updates the position of a user interface object. The GUI update unit 278 updates the GUI. For example, the GUI update unit 278 prepares display information and sends the display information to the graphic module 232 for display on the touch-sensitive display.
[0139] In some embodiments, the event processing unit(s) 290 includes or has access to the data update unit 276, the object update unit 277, and the GUI update unit 278. In some embodiments, the data update unit 276, the object update unit 277, and the GUI update unit 278 are included in a single module of the individual application 236-1 or the application view 291. In other embodiments, they are included in two or more software modules.
[0140] The foregoing description regarding event processing of a user's touch on the touch-sensitive display also applies to other forms of user input for operating the multifunctional device 200 using an input device, but it should be understood that not all of them are initiated on the touch screen. For example, the movement of the mouse and the pressing of the mouse button, the movement of contact such as tapping, dragging, and scrolling on the touch pad, pen stylus input, the movement of the device, voice commands, detected eye movement, biometric input, and / or any combination thereof, optionally in association with a single or multiple presses or holds of the keyboard, are used as input corresponding to sub-events that define events to be optionally recognized.
[0141] FIG. 3 shows a portable multifunctional device 200 having a touch screen 212, according to some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 300. In this embodiment, as well as other embodiments described below, a user can select one or more of those graphics by performing gestures on the graphics using, for example, one or more fingers 302 (not drawn to scale in the figure) or one or more styli 303 (not drawn to scale in the figure). In some embodiments, the selection of one or more graphics is performed when the user interrupts contact with the one or more graphics. In some embodiments, the gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, upward and / or downward), and / or rolling (from right to left, from left to right, upward and / or downward) of a finger in contact with the device 200. In some implementations or situations, an accidental contact with a graphic does not select the graphic. For example, if the gesture corresponding to selection is a tap, a swipe gesture that sweeps over an application icon does not optionally select the corresponding application.
[0142] Device 200 also includes one or more physical buttons, such as a "home" or menu button 304. As described above, menu button 304 is used to navigate to any application 236 within the set of applications running on device 200. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on touch screen 212.
[0143] In one embodiment, device 200 includes a touch screen 212, a menu button 304, a push button 306 for turning the device on / off and locking the device, volume adjustment button(s) 308, a subscriber identity module (SIM) card slot 310, a headset jack 312, and a docking / charging external port 224. The push button 306 is optionally used to turn the device on / off by pressing the button and holding it pressed for a predetermined period, to lock the device by pressing the button and releasing it before a predetermined time has elapsed, and / or to unlock the device or initiate an unlock process. In an alternative embodiment, device 200 also accepts verbal input via a microphone 213 to activate or deactivate some functions. Device 200 optionally also includes one or more contact intensity sensors 265 for detecting the intensity of contact on the touch screen 212 and / or one or more haptic output generators 267 for generating haptic output to the user of device 200.
[0144] FIG. 4 is a block diagram of an exemplary multifunctional device having a display and a touch sensing surface, according to some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a child's learning toy), gaming system, or control device (e.g., a home or industrial controller). Device 400 typically includes one or more processing units (CPUs) 410, one or more networks or other communication interfaces 460, memory 470, and one or more communication buses 420 interconnecting these components. Communication bus 420 optionally includes circuitry (sometimes called a chipset) that interconnects and controls communication between system components. Device 400 includes an input / output (I / O) interface 430 that includes a display 440, which is typically a touch screen display. I / O interface 430 also optionally includes a keyboard and / or mouse (or other pointing device) 450, as well as a touch pad 455, a touch output generator 457 (similar to, e.g., touch output generator(s) 267 described above with reference to FIG. 2A) for generating a touch output on device 400, and a sensor 459 (e.g., an optical sensor, an acceleration sensor, a proximity sensor, a touch sensing sensor, and / or a contact intensity sensor similar to contact intensity sensor(s) 265 described above with reference to FIG. 2A). Memory 470 includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 470 optionally includes one or more storage devices located remotely from the CPU(s) 410.In some embodiments, memory 470 stores programs, modules, and data structures, or subsets thereof, similar to the programs, modules, and data structures stored in memory 202 of portable multifunction device 200 (FIG. 2A). Additionally, memory 470 optionally stores additional programs, modules, and data structures not present in memory 202 of portable multifunction device 200. For example, memory 470 of device 400 optionally stores a drawing module 480, a presentation module 482, a word processing module 484, a website creation module 486, a disk authoring module 488, and / or a spreadsheet module 490, while memory 202 of portable multifunction device 200 (FIG. 2A) optionally does not store these modules.
[0145] Each element identified above in FIG. 4 is stored in one or more of the memory devices mentioned above in some embodiments. Each of the modules identified above corresponds to a set of instructions for performing the functions described above. The modules or programs (e.g., sets of instructions) identified above need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise reconfigured in various embodiments. In some embodiments, memory 470 stores a subset of the modules and data structures identified above. Additionally, memory 470 stores additional modules and data structures not described above.
[0146] Here, for example, attention is drawn to embodiments of a user interface that can be implemented on portable multifunction device 200.
[0147] FIG. 5A shows an exemplary user interface for a menu of applications on a portable multifunctional device 200, according to some embodiments. A similar user interface is implemented on device 400. In some embodiments, the user interface 500 includes the following elements, or a subset or superset thereof.
[0148] Signal strength indicator(s) 502 for wireless communication(s) such as cellular and Wi-Fi signals, ● Time 504, ● Bluetooth indicator 505, ● Battery status indicator 506, ● A tray 508 having icons of frequently used applications such as ○ An icon 516 of the phone module 238 labeled "Phone", optionally including an indicator 514 of the number of missed calls or voicemail messages, ○ An icon 518 of the email client module 240 labeled "Mail", optionally including an indicator 510 of the number of unread emails, ○ An icon 520 of the browser module 247 labeled "Browser", and ○ An icon 522 for the video and music player module 252, also referred to as the iPod (trademark of Apple Inc.) module 252, labeled "iPod", and ● Icons of other applications such as ○ An icon 524 of the IM module 241 labeled "Message", ○ An icon 526 of the calendar module 248 labeled "Calendar", ○ An icon 528 of the image management module 244 labeled "Photos", ○ An icon 530 of the camera module 243 labeled "Camera", ○ An icon 532 of the online video module 255 labeled "Online Video", ○ An icon 534 of the stock price widget 249-2 labeled with "Stock Price", ○ An icon 536 of the map module 254 labeled with "Map", ○ An icon 538 of the weather widget 249-1 labeled with "Weather", ○ An icon 540 of the alarm clock widget 249-4 labeled with "Clock", ○ An icon 542 of the training support module 242 labeled with "Training Support", ○ An icon 544 of the memo module 253 labeled with "Memo", and ○ An icon 546 of the settings application or module labeled with "Settings" that provides access to the settings of the device 200 and its various applications 236.
[0149] Note that the labels of the icons shown in FIG. 5A are merely illustrative. For example, the icon 522 for the video and music player module 252 may optionally be labeled with "Music" or "Music Player". Other labels may optionally be used for the various application icons. In some embodiments, the label for an individual application icon includes the name of the application corresponding to the individual application icon. In some embodiments, the label for a particular application icon is different from the name of the application corresponding to that particular application icon.
[0150] FIG. 5B shows an exemplary user interface on a device (e.g., device 400 of FIG. 4) having a touch sensing surface 551 (e.g., tablet or touch pad 455 of FIG. 4) separate from a display 550 (e.g., touch screen display 212). The device 400 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 459) for detecting the intensity of a contact on the touch sensing surface 551, and / or one or more haptic output generators 457 for generating haptic output to the user of the device 400.
[0151] Some of the following examples are described with reference to inputs on touch screen display 212 (when the touch sensing surface and the display are combined), but in some embodiments, the device detects inputs on a touch sensing surface separate from the display, as shown in FIG. 5B. In some embodiments, this touch sensing surface (e.g., 551 in FIG. 5B) has a major axis (e.g., 552 in FIG. 5B) corresponding to the major axis (e.g., 553 in FIG. 5B) on the display (e.g., 550). According to these embodiments, the device detects contacts (e.g., 560 and 562 in FIG. 5B) with the touch sensing surface 551 at locations corresponding to respective locations on the display (e.g., in FIG. 5B, 560 corresponds to 568 and 562 corresponds to 570). In this way, when the touch sensing surface is separate from the display, user inputs (e.g., contacts 560 and 562 and their movements) detected by the device on the touch sensing surface (e.g., 551 in FIG. 5B) are used by the device to operate the user interface on the display (e.g., 550 in FIG. 5B) of the multifunctional device. It should be understood that a similar method is optionally used for other user interfaces described herein.
[0152] In addition, while the following examples are given primarily with reference to finger inputs (e.g., finger contact, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs may be replaced by inputs from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be a mouse click (e.g., instead of a contact) followed by a movement of the cursor along the path of the swipe (e.g., instead of a movement of the contact). As another example, a tap gesture may optionally be a mouse click while the cursor is located over the location of the tap gesture (e.g., instead of detecting a contact and then ceasing to detect the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice may optionally be used simultaneously, or that mouse and finger contacts may optionally be used simultaneously.
[0153] FIG. 6A shows an exemplary personal electronic device 600. The device 600 includes a body 602. In some embodiments, the device 600 includes some or all of the features described in connection with devices 200 and 400 (e.g., FIGS. 2A-4). In some embodiments, the device 600 has a touch-sensitive display screen 604, hereinafter referred to as a touch screen 604. Alternatively, or in addition to the touch screen 604, the device 600 has a display and a touch-sensitive surface. Together with devices 200 and 400, in some embodiments, the touch screen 604 (or touch-sensitive surface) has one or more intensity sensors that detect the intensity of an applied contact (e.g., a touch). One or more intensity sensors of the touch screen 604 (or touch-sensitive surface) provide output data representing the intensity of the touch. The user interface of the device 600 responds to the touch based on the intensity of the touch, which means that touches of different intensities can call different user interface operations on the device 600.
[0154] Techniques for detecting and processing touch intensity can be found, for example, in the related applications: "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application" of International Patent Application No. PCT / US2013 / 040061 filed on May 8, 2013, and "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships" of International Patent Application No. PCT / US2013 / 069483 filed on November 11, 2013, each of which is hereby incorporated by reference in its entirety.
[0155] In some embodiments, device 600 has one or more input mechanisms 606 and 608. Input mechanisms 606 and 608, if included, are physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 600 has one or more attachment mechanisms. Such attachment mechanisms, if included, can enable device 600 to be attached, for example, to hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms enable the user to wear device 600.
[0156] FIG. 6B shows an exemplary personal electronic device 600. In some embodiments, device 600 includes some or all of the components described in connection with FIGS. 2A, 2B, and 4. Device 600 has a bus 612 that operably couples I / O section 614 to one or more computer processors 616 and memory 618. I / O section 614 is connected to a display 604 that may have a touch sensing component 622 and optionally a touch intensity sensing component 624. In addition, I / O section 614 is connected to a communication unit 630 that receives application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication technologies. Device 600 includes input mechanisms 606 and / or 608. Input mechanism 606 is, for example, a rotatable input device or a pressable and rotatable input device. Input mechanism 608 is a button in some embodiments.
[0157] Input mechanism 608 is a microphone in some embodiments. Personal electronic device 600 includes various sensors such as, for example, a GPS sensor 632, an accelerometer 634, a direction sensor 640 (e.g., a compass), a gyroscope 636, a motion sensor 638, and / or combinations thereof, all of which are operably connected to I / O section 614.
[0158] The memory 618 of the personal electronic device 600 is a non-transitory computer-readable storage medium that stores computer-executable instructions. For example, when executed by one or more computer processors 616, the computer processors are caused to perform the following techniques and processes. Those computer-executable instructions may also be stored and / or transmitted in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as, for example, a computer-based system, a system including a processor, or another system capable of fetching and executing instructions from an instruction execution system, apparatus, or device. The personal electronic device 600 is not limited to the components and configurations of FIG. 6B and may include other components or additional components in multiple configurations.
[0159] As used herein, the term "affordance" refers to, for example, user-interactive graphical user interface objects displayed on the display screens of devices 200, 400, 600, and / or 900 (FIGS. 2A, 4, 6A-6B, 9, and 10A-10B). For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each constitute an affordance.
[0160] As used herein, the term "focus selector" refers to an input element that indicates the current part of the user interface with which the user is interacting. In some implementations that include a cursor or other location marker, the cursor serves as the "focus selector" such that while the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), an input (e.g., a press input) detected on a touch-sensitive surface (e.g., touchpad 455 in FIG. 4, or touch-sensitive surface 551 in FIG. 5B) causes that particular user interface element to be adjusted according to the detected input. In some implementations that include a touch screen display (e.g., touch-sensitive display system 212 in FIG. 2A, or touch screen 212 in FIG. 5A) that enables direct interaction with user interface elements on the touch screen display, a contact detected on the touch screen serves as the "focus selector" such that an input (e.g., a press input by the contact) detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touch screen display causes that particular user interface element to be adjusted according to the detected input. In some implementations, the focus is moved from one region of the user interface to another region of the user interface without moving the corresponding cursor or contact on the touch screen display (e.g., by using the tab key or arrow keys to move the focus from one button to another), and in these implementations, the focus selector moves in accordance with the movement of the focus between different regions of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or a contact on a touch screen display) that is controlled by the user to communicate the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface through which the user intends to interact).For example, the location of a focus selector (e.g., a cursor, contact, or selection box) over an individual button while a press input is detected on a touch sensing surface (e.g., a touch pad or touch screen) indicates that the user intends to activate that individual button (as opposed to other user interface elements shown on the device's display).
[0161] As used in this specification and the claims, the term "characteristic strength" of a contact refers to the characteristics of that contact based on one or more strengths of the contact. In some embodiments, the characteristic strength is based on a plurality of strength samples. The characteristic strength is optionally based on a set of strength samples collected during a predetermined time (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined number of strength samples, i.e., a predetermined event (e.g., after detecting a contact, before detecting lift-off of the contact, before or after detecting the start of movement of the contact, before detecting the end of the contact, before or after detecting an increase in the strength of the contact, and / or before or after detecting a decrease in the strength of the contact). The characteristic strength of a contact is optionally based on one or more of the maximum value of the strength of the contact, the median value of the strength of the contact, the average value of the strength of the contact, the top 10 percent value of the strength of the contact, the value that is half of the maximum value of the strength of the contact, the value that is 90 percent of the maximum value of the strength of the contact, etc. In some embodiments, the duration of the contact is used when determining the characteristic strength (e.g., when the characteristic strength is the average of the strength of the contact over time). In some embodiments, the characteristic strength is compared to a set of one or more strength thresholds to determine whether an operation has been performed by a user. For example, the set of one or more strength thresholds includes a first strength threshold and a second strength threshold. In this example, a contact having a characteristic strength that does not exceed the first threshold results in a first operation, a contact having a characteristic strength that exceeds the first strength threshold but does not exceed the second strength threshold results in a second operation, and a contact having a characteristic strength that exceeds the second threshold results in a third operation. In some embodiments, the comparison of the characteristic strength to one or more thresholds is not used to determine which of a first or second operation to perform, but rather is used to determine whether to perform one or more operations (e.g., perform an individual operation or cancel the execution of an individual operation).
[0162] In some embodiments, for the purpose of determining characteristic intensity, a portion of a gesture is specified. For example, the touch sensing surface receives a continuous swipe contact that transitions from a start location point where the contact intensity increases to an end location point. In this example, the characteristic intensity of the contact at the end location is based on only a portion of the continuous swipe contact, rather than the entire swipe contact (e.g., only a portion of the swipe contact at the end location). In some embodiments, a smoothing algorithm is applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of a non-weighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or dips in the width of the swipe contact intensity for the purpose of determining characteristic intensity.
[0163] The intensity of a contact on the touch sensing surface is characterized relative to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold typically corresponds to the intensity at which the device performs an operation associated with clicking a button or trackpad of a physical mouse. In some embodiments, the deep press intensity threshold typically corresponds to the intensity at which the device performs an operation different from an operation associated with clicking a button or trackpad of a physical mouse. In some embodiments, when a contact is detected that has a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact detection intensity threshold below which contact is not detected), the device moves the focus selector in accordance with the movement of the contact on the touch sensing surface without performing an operation associated with the light press intensity threshold or the deep press intensity threshold. Generally, unless otherwise specified, these intensity thresholds are consistent among different sets of user interface values.
[0164] An increase in the characteristic strength of a contact from a strength below a light pressing strength threshold to a strength between the light pressing strength threshold and a deep pressing strength threshold may be referred to as an input of "light pressing". An increase in the characteristic strength of a contact from a strength below a deep pressing strength threshold to a strength above the deep pressing strength threshold may be referred to as an input of "deep pressing". An increase in the characteristic strength of a contact from a strength below a contact detection strength threshold to a strength between the contact detection strength threshold and the light pressing strength threshold may be referred to as a detection of a contact on the touch surface. A decrease in the characteristic strength of a contact from a strength above the contact detection strength threshold to a strength below the contact detection strength threshold may be referred to as a detection of a lift-off of the contact from the touch surface. In some embodiments, the contact detection strength threshold is zero. In some embodiments, the contact detection strength threshold is greater than zero.
[0165] In some embodiments described herein, in response to detecting a gesture that includes an individual pressing input, or in response to detecting an individual pressing input performed by an individual contact (or contacts), one or more operations are performed, and the individual pressing input is detected at least in part based on detecting an increase in the strength of a contact (or contacts) above a pressing input strength threshold. In some embodiments, the individual operation is performed in response to detecting an increase in the strength of an individual contact above the pressing input strength threshold (e.g., the "downstroke" of an individual pressing input). In some embodiments, the pressing input includes an increase in the strength of an individual contact above the pressing input strength threshold and a subsequent decrease in the strength of the contact below the pressing input strength threshold, and the individual operation is performed in response to detecting a subsequent decrease in the strength of the individual contact below the pressing input threshold (e.g., the "upstroke" of an individual pressing input).
[0166] In some embodiments, the device employs intensity hysteresis to avoid spurious inputs sometimes referred to as "jitter", and the device defines or selects a hysteresis intensity threshold having a predefined relationship to the press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some other reasonable percentage of the press input intensity threshold). Thus, in some embodiments, a press input includes an increase in the intensity of an individual contact above the press input intensity threshold, and a subsequent decrease in the intensity of the contact below the hysteresis intensity threshold corresponding to the press input intensity threshold, and an individual operation is performed in response to detecting a subsequent decrease in the intensity of an individual contact below the hysteresis intensity threshold (e.g., the "upstroke" of an individual press input). Similarly, in some embodiments, a press input is detected only when the device detects an increase in the intensity of a contact from below the hysteresis intensity threshold to above the press input intensity threshold, and optionally, a subsequent decrease in the intensity of the contact to below the hysteresis intensity, and an individual operation is performed in response to detecting the press input (e.g., an increase in the intensity of the contact or a decrease in the intensity of the contact, depending on the situation).
[0167] For ease of explanation, the description of an operation performed in response to a press input associated with a press input intensity threshold, or a gesture including a press input, is optionally triggered in response to detecting any of an increase in the intensity of a contact above the press input intensity threshold, an increase in the intensity of a contact from below the hysteresis intensity threshold to above the press input intensity threshold, a decrease in the intensity of a contact below the press input intensity threshold, and / or a decrease in the intensity of a contact below the hysteresis intensity threshold corresponding to the press input intensity threshold. Further, in examples where an operation is described as being performed in response to detecting a decrease in the intensity of a contact below the press input intensity threshold, the operation is optionally performed in response to detecting a decrease in the intensity of a contact below a hysteresis intensity threshold corresponding to and lower than the press input intensity threshold. 3. Digital Assistant System
[0168] FIG. 7A shows a block diagram of a digital assistant system 700 according to various embodiments. In some embodiments, the digital assistant system 700 is implemented on a stand-alone computer system. In some embodiments, the digital assistant system 700 is distributed across multiple computers. In some embodiments, some of the modules and functions of the digital assistant are allocated to a server portion and a client portion, and the client portion resides on one or more user devices (e.g., device 104, 122, 200, 400, 600, or 900) as shown in FIG. 1, for example, and communicates with the server portion (e.g., server system 108) through one or more networks. In some embodiments, the digital assistant system 700 is an implementation of the server system 108 (and / or DA server 106) shown in FIG. 1. The digital assistant system 700 is only one example of a digital assistant system, and it should be noted that the digital assistant system 700 may have more or fewer components than those shown, may combine two or more components, or may have different configurations or arrangements of those components. The various components shown in FIG. 7A are implemented as hardware including one or more signal processing circuits and / or application specific integrated circuits, software instructions executed by one or more processors, firmware, or a combination thereof.
[0169] The digital assistant system 700 includes a memory 702, one or more processors 704, an input / output (I / O) interface 706, and a network communication interface 708. These components can communicate with each other via one or more communication buses or signal lines 710.
[0170] In some embodiments, the memory 702 includes a non-transitory computer-readable medium, such as a high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).
[0171] In some embodiments, the I / O interface 706 couples the input / output devices 716 of the digital assistant system 700, such as a display, keyboard, touch screen, and microphone, to the user interface module 722. The I / O interface 706, in conjunction with the user interface module 722, receives user inputs (e.g., voice input, keyboard input, touch input, etc.) and processes them as appropriate. In some embodiments, for example, when the digital assistant is implemented on a stand-alone user device, the digital assistant system 700 includes any of the components and I / O communication interfaces described with respect to the devices 200, 400, 600, or 900 of FIGS. 2A, 4, 6A-6B, 9, and 10A-10B. In some embodiments, the digital assistant system 700 represents the server portion of the digital assistant implementation and can interact with the user through a client-side portion resident on a user device (e.g., device 104, 200, 400, 600, or 900).
[0172] In some embodiments, network communication interface 708 includes one or more wired communication ports 712 and / or wireless transceiver circuitry 714. The wired communication port(s) transmit and receive communication signals via one or more wired interfaces such as Ethernet, Universal Serial Bus (USB), FireWire, etc. The wireless circuitry 714 transmits and receives RF signals and / or optical signals between the communication network and other communication devices. The wireless communication uses any one of a plurality of communication standards, communication protocols, and communication technologies such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. The network communication interface 708 enables communication between the digital assistant system 700 and a network such as the Internet, an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN), and between other devices.
[0173] In some embodiments, the memory 702, or the computer-readable storage medium of the memory 702, stores programs, modules, instructions, and data structures including all or a subset of the operating system 718, communication module 720, user interface module 722, one or more applications 724, and digital assistant module 726. In particular, the memory 702, or the computer-readable storage medium of the memory 702, stores instructions for executing the processes described below. One or more processors 704 execute these programs, modules, and instructions and read from / write to the data structures.
[0174] The operating system 718 (e.g., an embedded operating system such as Darwin, RTXC, LINUX, UNIX, iOS, OS X, WINDOWS, or VxWorks) includes various software components and / or drivers for controlling and managing common system tasks (e.g., memory management, storage device control, power management, etc.), and facilitates communication among various hardware, firmware, and software components.
[0175] The communication module 720 facilitates communication between the digital assistant system 700 and other devices via the network communication interface 708. For example, the communication module 720 communicates with the RF circuits 208 of electronic devices such as the devices 200, 400, and 600 shown in FIGS. 2A, 4, and 6A - 6B, respectively. The communication module 720 also includes various components for processing data received by the wireless circuit 714 and / or the wired communication port 712.
[0176] The user interface module 722 receives commands and / or inputs from the user via the I / O interface 706 (e.g., from a keyboard, touch screen, pointing device, controller, and / or microphone) and generates user interface objects on the display. The user interface module 722 also prepares and delivers outputs to the user via the I / O interface 706 (e.g., through a display, audio channel, speaker, touch pad, etc.) (e.g., speech, sound, animation, text, icons, vibration, tactile feedback, light, etc.).
[0177] Application 724 includes programs and / or modules configured to be executed by one or more processors 704. For example, when the digital assistant system is implemented on a stand-alone user device, Application 724 includes user applications such as games, calendar applications, navigation applications, or email applications. When the digital assistant system 700 is implemented on a server, Application 724 includes, for example, resource management applications, diagnostic applications, or scheduling applications.
[0178] Memory 702 also stores the digital assistant module 726 (or the server part of the digital assistant). In some embodiments, the digital assistant module 726 includes the following sub-modules, or subsets or supersets thereof: input / output processing module 728, speech-to-text (STT) processing module 730, natural language processing module 732, dialog flow processing module 734, task flow processing module 736, service processing module 738, and speech synthesis processing module 740. Each of these modules has access to one or more of the following systems or data and models of the digital assistant module 726, or subsets or supersets thereof: ontology 760, vocabulary index 744, user data 748, task flow model 754, service model 756, and ASR system 758.
[0179] In some embodiments, by using the processing modules, data, and models implemented in the digital assistant module 726, the digital assistant can perform at least some of the following: convert speech input to text, identify the user's intent expressed in the natural language input received from the user, actively elicit and obtain the information necessary to fully infer the user's intent (e.g., by clarifying words, games, intents, etc.), determine a task flow to satisfy the inferred intent, and execute that task flow to satisfy the inferred intent.
[0180] In some embodiments, as shown in FIG. 7B, the I / O processing module 728 interacts with the user through the I / O device 716 of FIG. 7A to obtain user input (e.g., speech input) and provide a response to the user input (e.g., as a speech output), or interacts with a user device (e.g., device 104, 200, 400, or device 600) through the network communication interface 708 of FIG. 7A. The I / O processing module 728 optionally obtains context information associated with the user input from the user device, either with the user input or immediately after receiving the user input. The context information includes user-specific data, vocabulary, and / or preferences related to the user input. In some embodiments, the context information also includes the software state and hardware state of the user device at the time the user request is received, and / or information about the user's surrounding environment at the time the user request is received. In some embodiments, the I / O processing module 728 also sends supplementary questions to the user regarding the user request and receives answers from the user. When a user request is received by the I / O processing module 728 and the user request includes speech input, the I / O processing module 728 transfers the speech input to the STT processing module 730 (or, a speech recognizer) for speech-to-text conversion.
[0181] The STT processing module 730 includes one or more ASR systems 758. The one or more ASR systems 758 can process the speech input received via the I / O processing module 728 to generate a recognition result. Each ASR system 758 includes a front-end speech preprocessor. This front-end speech preprocessor extracts representative features from the speech input. For example, the front-end speech preprocessor extracts spectral features that characterize the speech input as a sequence of representative multi-dimensional vectors by performing a Fourier transform on the speech input. Further, each ASR system 758 includes one or more speech recognition models (e.g., an acoustic model and / or a language model) and implements one or more speech recognition engines. Examples of speech recognition models include hidden Markov models, mixture Gaussian models, deep neural network models, n-gram language models, and other statistical models. Examples of speech recognition engines include a dynamic time warping-based engine and a weighted finite-state transducer (WFST)-based engine. The one or more speech recognition models and the one or more speech recognition engines are used to process the extracted representative features of the front-end speech preprocessor to generate intermediate recognition results (e.g., phonemes, sequences of phonemes, sub-words) and finally a text recognition result (words, sequences of words, sequences of tokens). In some embodiments, the speech input is at least partially processed by a third-party service or on the user's device (e.g., device 104, 200, 400, or device 600) to generate a recognition result. When the STT processing module 730 generates a recognition result that includes a text string (e.g., a word, a sequence of words, or a sequence of tokens), the recognition result is passed to the natural language processing module 732 for intent inference. In some embodiments, the STT processing module 730 generates a plurality of text representation candidates for the speech input. Each text representation candidate is a sequence of words or tokens corresponding to the speech input. In some embodiments, each text representation candidate is associated with a speech recognition confidence score.Based on this speech recognition reliability score, the STT processing module 730 ranks the text representation candidates and provides the n best (e.g., n highest-ranked) text representation candidates (singular or plural) to the natural language processing module 732 for intent inference (n is a predetermined integer greater than zero). For example, in one embodiment, only the highest-ranked (n = 1) text representation candidate is passed to the natural language processing module 732 for intent inference. In another embodiment, the five highest-ranked (n = 5) text representation candidates are passed to the natural language processing module 732 for intent inference.
[0182] Further details regarding the speech-to-text processing are described in U.S. Utility Patent Application No. 13 / 236,942, filed on September 20, 2011, entitled "Consolidating Speech Recognition Results", which is incorporated herein by reference in its entirety.
[0183] In some embodiments, the STT processing module 730 includes a vocabulary of recognizable words and / or accesses that vocabulary via the phonetic conversion module 731. Each vocabulary word is associated with one or more pronunciation candidates for that word represented in speech recognition phonetic characters. Specifically, the vocabulary of recognizable words includes words associated with multiple pronunciation candidates. For example, the vocabulary includes
Number
[0184] In some embodiments, pronunciation candidates are ranked based on the generality of the pronunciation candidates. For example, pronunciation candidate
Number
Number
Number
Number
Number
Number
[0185] When an utterance input is received, the STT processing module 730 is used to determine the phonemes corresponding to the utterance input (e.g., using an acoustic model), and then an attempt is made to determine the word that matches the phonemes (e.g., using a language model). For example, the STT processing module 730 first determines a sequence of phonemes corresponding to a part of the utterance input
Number
[0186] In some embodiments, the STT processing module 730 uses approximate matching techniques to determine the words in the utterance. Therefore, for example, the STT processing module 730 determines that a sequence of phonemes
Number
[0187] The natural language processing module 732 (the "natural language processor") of the digital assistant obtains the n best text representation candidates (singular or plural) (the "word sequence(s)" or "token sequence(s)") generated by the STT processing module 730 and attempts to associate each of those text representation candidates with one or more "feasible intents" recognized by the digital assistant. A "feasible intent" (or "user intent") represents a task that can be performed by the digital assistant and may have an associated task flow implemented within the task flow model 754. This associated task flow is a series of programmed actions and steps that the digital assistant performs to execute that task. The scope of the digital assistant's capabilities is determined according to the number and variety of task flows implemented and stored within the task flow model 754, or, in other words, according to the number and variety of "feasible intents" recognized by that digital assistant. However, the effectiveness of the digital assistant is also determined according to the assistant's ability to infer the correct "feasible intent(s)" from the user request expressed in natural language.
[0188] In some embodiments, in addition to the sequence of words or tokens obtained from the STT processing module 730, the natural language processing module 732 also receives context information associated with the user request, for example, from the I / O processing module 728. The natural language processing module 732 optionally uses that context information to clarify, complement, and / or further define the information contained within the text representation candidates received from the STT processing module 730. Context information includes, for example, user preferences, the hardware and / or software state of the user device, sensor information collected before, during, or immediately after the user request, and previous interactions (e.g., dialogs) between the digital assistant and the user. As described herein, the context information is dynamic in some embodiments and changes according to time, location, the content of the dialog, and other factors.
[0189] In some embodiments, natural language processing is based on, for example, ontology 760. Ontology 760 is a hierarchical structure including a number of nodes, and each node represents a "feasible intention" or represents an "attribute" or other "attributes" related to one or more of the "feasible intentions". As described above, a "feasible intention" represents a task that a digital assistant can execute, that is, the task is "feasible" or can be a target of implementation. An "attribute" represents a parameter associated with a feasible intention or associated with a subordinate aspect of another attribute. The link between the feasible intention node and the attribute node in ontology 760 defines how the parameter represented by the attribute node is involved in the task represented by the feasible intention node.
[0190] In some embodiments, ontology 760 is composed of feasible intention nodes and attribute nodes. Within ontology 760, each feasible intention node is directly linked to one or more attribute nodes or is linked through one or more intermediate attribute nodes. Similarly, each attribute node is directly linked to one or more feasible intention nodes or is linked through one or more intermediate attribute nodes. For example, as shown in FIG. 7C, ontology 760 includes a "restaurant reservation" node (i.e., a feasible intention node). The attribute nodes "restaurant", "date / time" (for reservation), and "number of participants" are directly linked to the feasible intention node (i.e., the "restaurant reservation" node), respectively.
[0191] Furthermore, the attribute nodes "Cuisine", "Price Range", "Phone Number", and "Location" are child nodes of the attribute node "Restaurant", and are each linked to the "Restaurant Reservation" node (i.e., the actionable intent node) via the intermediate attribute node "Restaurant". As another example, as shown in FIG. 7C, ontology 760 also includes a "Reminder Setting" node (i.e., another actionable intent node). The attribute nodes "Date / Time" (for reminder setting) and "Theme" (for reminder) are linked to the "Reminder Setting" node respectively. Since the attribute node "Date / Time" is related to both the task of performing a restaurant reservation and the task of setting a reminder, the attribute node "Date / Time" is linked to both the "Restaurant Reservation" node and the "Reminder Setting" node within ontology 760.
[0192] An actionable intent node, along with its linked attribute nodes, is described as a "domain". In this discussion, each domain is associated with an individual actionable intent and refers to a group of nodes (and the relationships between those nodes) associated with that particular actionable intent. For example, ontology 760 shown in FIG. 7C includes an example of a restaurant reservation domain 762 and an example of a reminder domain 764 within ontology 760. The restaurant reservation domain includes the actionable intent node "Restaurant Reservation", the attribute nodes "Restaurant", "Date / Time", and "Number of Participants", and the child attribute nodes "Cuisine", "Price Range", "Phone Number", and "Location". The reminder domain 764 includes the actionable intent node "Reminder Setting", and the attribute nodes "Theme" and "Date / Time". In some embodiments, ontology 760 is composed of multiple domains. Each domain shares one or more attribute nodes with one or more other domains. For example, the "Date / Time" attribute node is associated with a number of different domains (such as a scheduling domain, a travel reservation domain, a movie ticket domain, etc.) in addition to the restaurant reservation domain 762 and the reminder domain 764.
[0193] FIG. 7C shows two exemplary domains within ontology 760, and other domains include, for example, "search for a movie", "initiate a phone call", "find a route", "schedule a meeting", "send a message", "provide an answer to a question", "read a list", "provide navigation instructions", and "provide instructions regarding a task", etc. The "send a message" domain is associated with the actionable intent of "send a message" and further includes attribute nodes such as "recipient(s)", "message type", and "message body". The attribute node "recipient" is further defined by sub-attribute nodes such as, for example, "recipient name" and "message address".
[0194] In some embodiments, ontology 760 includes all domains (and thus actionable intents) that a digital assistant can understand and perform. In some embodiments, ontology 760 is modified by adding or removing entire domains or nodes, or by modifying the relationships between nodes within ontology 760, etc.
[0195] In some embodiments, nodes associated with multiple related actionable intents are clustered under a "superordinate domain" within ontology 760. For example, the "travel" superordinate domain includes a cluster of attribute nodes and actionable intent nodes related to travel. Actionable intent nodes related to travel include "airline reservation", "hotel reservation", "car rental", "know the way", "find interesting points", etc. Actionable intent nodes under the same superordinate domain (e.g., the "travel" superordinate domain) share a number of attribute nodes. For example, actionable intent nodes related to "flight reservation", "hotel reservation", "car rental", "know the route", and "find interesting places" share one or more of the attribute nodes "departure location", "destination", "departure date / time", "arrival date / time", and "number of participants".
[0196] In some embodiments, each node within ontology 760 is associated with a set of words and / or phrases related to the attributes or actionable intents represented by that node. The individual set of words and / or phrases associated with each node is the so-called "vocabulary" associated with that node. The individual set of words and / or phrases associated with each node is stored in vocabulary index 744 in relation to the attributes or actionable intents represented by that node. For example, returning to FIG. 7B, words such as "food", "drink", "dish", "hunger", "eat", "pizza", "fast food", "meal" are included in the vocabulary associated with the node related to the attribute of "restaurant". As another example, words and phrases such as "call", "phone", "dial", "ring", "call this number", "make a call to" are included in the vocabulary associated with the node related to the actionable intent of "initiate a phone call". Vocabulary index 744 optionally includes words and phrases in different languages.
[0197] The natural language processing module 732 receives a candidate text representation (e.g., a string (singular or plural) or a token sequence (singular or plural)) from the STT processing module 730, and for each candidate representation, determines which node the words in the candidate character representation imply. In some embodiments, if a word or phrase within the text representation candidate is found (via the vocabulary index 744) to be associated with one or more nodes within the ontology 760, then that word or phrase "triggers" or "activates" those nodes. Based on the amount and / or relative importance of the activated nodes, the natural language processing module 732 selects, as the task that the user intends the digital assistant to perform, one of those possible intents. In some embodiments, the domain having the most "triggered" nodes is selected. In some embodiments, the domain having the highest confidence value (e.g., based on the relative importance of the various nodes that are triggered) is selected. In some embodiments, the domain is selected based on a combination of the number and importance of the triggered nodes. In some embodiments, additional factors, such as whether the digital assistant has accurately interpreted similar requests from the user in the past, are also considered when selecting the nodes.
[0198] The user data 748 includes user-specific information such as the user's unique vocabulary, user preferences, user address, the user's default and second languages, the user's contact list, and other short-term or long-term information regarding each user. In some embodiments, the natural language processing module 732 uses this user-specific information to complement the information contained within the user input to further define the user intent. For example, with respect to the user request "invite my friends to my birthday party", the natural language processing module 732 can access the user data 748 without asking the user to explicitly provide such information within the user request, in order to determine who the "friends" are and when and where the "birthday party" will be held.
[0199] In some embodiments, it should be recognized that the natural language processing module 732 is implemented using one or more machine learning mechanisms (e.g., neural networks). Specifically, the one or more machine learning mechanisms are configured to receive text representation candidates and context information associated with the text representation candidates. Based on the text representation candidates and the associated context information, the one or more machine learning mechanisms are configured to determine intent reliability scores across a set of actionable intent candidates. The natural language processing module 732 can select one or more actionable intent candidates from the set of actionable intent candidates based on the determined intent reliability scores. In some embodiments, an ontology (e.g., ontology 760) is also used to select one or more actionable intent candidates from the set of actionable intent candidates.
[0200] Other details of the ontology search based on the token string are described in U.S. Utility Patent Application No. 12 / 341,743, filed Dec. 22, 2008, entitled "Method and Apparatus for Searching Using an Active Ontology", which is hereby incorporated by reference in its entirety.
[0201] In some embodiments, when the natural language processing module 732 identifies an intent (or domain) that can be implemented based on a user request, the natural language processing module 732 generates a structured query to represent the identified implementable intent. In some embodiments, this structured query includes parameters for one or more nodes within the domain related to the implementable intent, and at least some of these parameters are input with specific information and requirements specified within the user request. For example, the user says "Make me a dinner reservation at a sushi place at 7". In this case, the natural language processing module 732 can accurately identify that the implementable intent based on the user input is "restaurant reservation". According to the ontology, the structured query related to the "restaurant reservation" domain includes parameters such as {dish}, {time}, {date}, {number of participants}, etc. In some embodiments, based on the utterance input and the text derived from the utterance input using the STT processing module 730, the natural language processing module 732 generates a partial structured query related to the restaurant reservation domain, and this partial structured query includes the parameter {dish = "sushi"} and the parameter {time = "7 pm"}. However, in this example, the information contained in the user's utterance is insufficient to complete the structured query associated with the domain. Therefore, other necessary parameters such as {number of participants} and {date} are not specified in the structured query based on the currently available information. In some embodiments, the natural language processing module 732 appends the received context information as input to some of the parameters of this structured query. For example, in some embodiments, when the user requests a "nearby" sushi restaurant, the natural language processing module 732 appends the GPS coordinates from the user device as input to the {location} parameter within the structured query.
[0202] In some embodiments, the natural language processing module 732 identifies a plurality of possible intent candidates for each text representation candidate received from the STT processing module 730. Further, in some embodiments, for each of the identified possible intent candidates, an individual (partial or complete) structured query is generated. The natural language processing module 732 determines an intent confidence score for each of the possible intent candidates and ranks those possible intent candidates based on that intent confidence score. In some embodiments, the natural language processing module 732 passes the generated structured query(ies), including any input parameters, to the task flow processing module 736 (the "task flow processor"). In some embodiments, the structured query(ies) for the m best (e.g., the m highest ranked) possible intent candidates are provided to the task flow processing module 736 (where m is a predetermined integer greater than zero). In some embodiments, the structured query(ies) for the m best possible intent candidates are provided to the task flow processing module 736 along with the corresponding text representation candidate(s).
[0203] Other details of the inference of user intent based on a plurality of possible intent candidates determined from a plurality of text representation candidates of the utterance input are described in U.S. Patent Application No. 14 / 298,725, filed Jun. 6, 2014, entitled "System and Method for Inferring User Intent From Speech Inputs", which is hereby incorporated by reference in its entirety.
[0204] The task flow processing module 736 is configured to receive structured query(ies) from the natural language processing module 732, complete the structured query as necessary, and execute actions required to "complete" the user's ultimate request. In some embodiments, various procedures required to complete these tasks are provided within the task flow model 754. In some embodiments, the task flow model 754 includes procedures for obtaining additional information from the user and a task flow for executing actions associated with actionable intents.
[0205] As described above, in order to complete the structured query, the task flow processing module 736 needs to start an additional dialogue with the user to obtain additional information and / or remove the ambiguity of potentially ambiguous statements. If such an interaction is necessary, the task flow processing module 736 calls the dialogue flow processing module 734 to engage in a dialogue with the user. In some embodiments, the dialogue flow processing module 734 determines how (and / or when) to request additional information from the user, and receives and processes the user response. Through the I / O processing module 728, questions are provided to the user and answers are received from the user. In some embodiments, the dialogue flow processing module 734 presents dialogue output to the user via audio and / or visual output, and receives input from the user via verbal or physical (e.g., click) responses. Continuing with the above embodiment, when the task flow processing module 736 calls the dialogue flow processing module 734 to determine the "number of participants" and "date" information for a structured query associated with the domain "restaurant reservation", the dialogue flow processing module 734 generates questions such as "For how many people?" and "On which day?" and passes them to the user. When an answer is received from the user, the dialogue flow processing module 734 then either adds the missing information to the structured query or passes that information to the task flow processing module 736 to complete the missing information from the structured query.
[0206] When the task flow processing module 736 completes a structured query regarding an actionable intent, the task flow processing module 736 proceeds to execute the final task associated with that actionable intent. Thus, the task flow processing module 736 executes steps and instructions within the task flow model according to specific parameters included within the structured query. For example, a task flow model regarding the actionable intent of "restaurant reservation" includes steps and instructions to contact the restaurant and actually request a reservation for a specific number of participants at a specific time. For example, using a structured query such as {restaurant reservation, restaurant = ABC Cafe, date = 3 / 12 / 2012, time = 7:00 PM, number of participants = 5}, the task flow processing module 736 (1) logs on to the server of ABC Cafe or a restaurant reservation system such as OPENTABLE (registered trademark), (2) enters the information of date, time, and number of participants into a form on the website, (3) submits that form, and (4) enters a calendar item regarding that reservation into the user's calendar.
[0207] In some embodiments, the task flow processing module 736 employs the assistance of a service processing module 738 (the "service processing module") to complete a task requested by a user input or to provide an answer to information requested by a user input. For example, instead of the task flow processing module 736, the service processing module 738 makes a phone call, sets a calendar item, invokes a map search, invokes or interacts with other user applications installed on the user device, or invokes or interacts with a third-party service (such as a restaurant reservation portal, a social networking website, a banking portal, etc.). In some embodiments, the protocols and application programming interfaces (APIs) required by each service are specified by individual service models within the service model 756. The service processing module 738 accesses an appropriate service model for the service and generates requests for the service in accordance with the protocols and APIs required by the service according to that service model.
[0208] For example, if a restaurant supports an online reservation service, the restaurant submits a service model that specifies the parameters necessary to make a reservation and the API for communicating the values of those necessary parameters to the online reservation service. When requested by the task flow processing module 736, the service processing module 738 uses the web address stored within the service model to establish a network connection with the online reservation service and transmits the necessary reservation parameters (e.g., time, date, number of participants) in a format compliant with the API of the online reservation service to the online reservation interface.
[0209] In some embodiments, the natural language processing module 732, the dialog flow processing module 734, and the task flow processing module 736 are used collectively and iteratively to infer and define the user's intent, obtain information for further clarifying and narrowing down the user intent, and ultimately generate a response (i.e., an output to the user or completion of a task) to satisfy the user's intent. The generated response is a dialog response to the utterance input that at least partially satisfies the user's intent. Further, in some embodiments, the generated response is output as an utterance output. In these embodiments, the generated response can be sent to an utterance synthesis processing module 740 (e.g., an utterance synthesizer), and the utterance synthesis processing module can process it to synthesize an utterance-formatted dialog response. In still other embodiments, the generated response is data content related to satisfying the user request within the utterance input.
[0210] In an example where the task flow processing module 736 receives a plurality of structured queries from the natural language processing module 732, the task flow processing module 736 first processes the first structured query among the received structured queries to complete the first structured query and / or attempts to execute one or more tasks or actions represented by the first structured query. In some examples, the first structured query corresponds to the highest ranked actionable intent. In other examples, the first structured query is selected from the received structured queries based on a combination of the corresponding speech recognition confidence score and the corresponding intent confidence score. In some examples, if the task flow processing module 736 encounters an error during the processing of the first structured query (e.g., due to a necessary parameter being undetermined), the task flow processing module 736 can proceed to select and process a second structured query among the received structured queries that corresponds to a lower ranked actionable intent. This second structured query is selected based on, for example, the speech recognition confidence score of the corresponding text representation candidate, the intent confidence score of the corresponding actionable intent candidate, the missing necessary parameter within the first structured query, or any combination thereof.
[0211] The speech synthesis processing module 740 is configured to synthesize a speech output for presentation to the user. The speech synthesis processing module 740 synthesizes the speech output based on the text provided by the digital assistant. For example, the generated dialog response is in the form of a text string. The speech synthesis processing module 740 converts the text string into an audible speech output. The speech synthesis processing module 740 uses any suitable speech synthesis technique, including but not limited to, waveform concatenation synthesis, unit selection synthesis, diphone synthesis, domain-limited synthesis, formant synthesis, prosody synthesis, hidden Markov model (HMM)-based synthesis, and sine wave synthesis, to generate the speech output from the text. In some embodiments, the speech synthesis processing module 740 is configured to synthesize individual words based on a sequence of phonemes corresponding to the words. For example, a sequence of phonemes is associated with the words in the generated dialog response. The sequence of phonemes is stored in the metadata associated with the words. The speech synthesis processing module 740 is configured to directly process the sequence of phonemes in the metadata to synthesize the words in voice form.
[0212] In some embodiments, instead of (or in addition to) using the speech synthesis processing module 740, speech synthesis is performed on a remote device (e.g., the server system 108), and the synthesized speech is transmitted to the user device for output to the user. For example, this can be implemented in some implementations where the output for the digital assistant is generated on the server system. Also, since the server system generally has more processing power or resources than the user device, it is possible to obtain a higher quality speech output than the output that would be practical in the case of client-side synthesis.
[0213] For more details about the digital assistant, reference may be made to U.S. Utility Application No. 12 / 987,982, entitled "Intelligent Automated Assistant", filed on January 10, 2011, and U.S. Utility Application No. 13 / 251,088, entitled "Generating and Processing Task Items That Represent Tasks to Perform", filed on September 30, 2011, the entire disclosures of which are incorporated herein by reference. 4. Determining whether the user's visual attention is directed towards the electronic device while the user is speaking
[0214] FIG. 8A shows a system 800 for determining whether the user's visual attention is directed towards an electronic device (e.g., device 104, 122, 200, 400, 600, or 900) while the user is speaking, according to various examples. In some embodiments, the system 800 is implemented on a stand-alone computer system (e.g., device 104, 122, 200, 400, 600, or 900). In some embodiments, the system 800 is distributed across multiple computers. For example, some of the components and functions of the system 800 are allocated to a server portion and a client portion, and the client portion resides on one or more user devices (e.g., device 104, 122, 200, 400, 600, or 900), as shown in FIG. 1 for example, and communicates with the server portion (e.g., server system 108) through one or more networks.
[0215] The system 800 is implemented using hardware, software, or a combination of hardware and software to perform the principles discussed herein. In some embodiments, the components and functions of the system 800 are implemented within the digital assistant module 726 as described above with reference to FIGS. 7A-7C. For example, each component of the system 800 is implemented as a set of computer-executable instructions stored in the memory 702.
[0216] System 800 is exemplary, and thus, System 800 can have more or fewer components than shown, can combine two or more components, or can have different configurations or arrangements of components. Further, the following discussion describes functionality performed by a single component of System 800, but it should be understood that such functionality can be performed by other components of System 800 and that such functionality can be performed by two or more components of System 800.
[0217] System 800 receives an audio stream and a video stream simultaneously. In some embodiments, one or more cameras of the device implementing System 800 (e.g., RGB camera(s), infrared (IR) camera(s), depth camera(s), etc.) capture the video stream. In some embodiments, an external camera(s) of the device implementing System 800 captures the video stream. Similarly, in some embodiments, one or more audio sensors of the device implementing System 800 (e.g., microphone(s)) sample the audio stream. In some embodiments, an external audio sensor(s) of the device implementing System 800 samples the audio stream. In some embodiments, the audio stream includes audio sampled by the audio sensor(s) of the device and audio sampled by the external audio sensor(s) of the device.
[0218] In some embodiments, system 800 includes a preprocessing module 802. The preprocessing module 802 is configured to process an audio stream and / or a video stream. As an example, the preprocessing module 802 processes the audio stream using voice emphasis techniques and / or noise reduction techniques. As another example, the preprocessing module 802 identifies (individual or respective) portions (e.g., cropped portions) of the video stream so as to include different user(s). For example, the preprocessing module 802 tracks the head position(s) of the user(s) within the video stream. In some embodiments, identifying (individual or respective) portions of the video stream so as to include different user(s) includes cropping the video stream to obtain the (individual or respective) portions. For example, the first portion of the video stream includes the head (e.g., face) of the first user, the second portion of the video stream includes the head (e.g., face) of the second user, and so on.
[0219] In some embodiments, system 800 includes a memory buffer 804. Memory buffer 804 is configured to store an (optionally pre - processed) audio stream and an (optionally pre - processed) video stream. In an example where each portion (e.g., a cropped portion) of the video data stream is identified as including a different user, system 800 includes separate instances of memory buffer 804 for each user. For example, a first instance of memory buffer 804 includes an audio stream and a first portion of the video stream that includes the head (e.g., face) of a first user, and a second instance of memory buffer 804 includes an audio stream and a second portion of the video stream that includes the head (e.g., face) of a second user. As described below, the audiovisual call module 806 can separately process the stored content of each instance of memory buffer 804 to determine, for each user, whether the user's visual attention is directed towards the electronic device while the user is speaking.
[0220] In some embodiments, memory buffer 804 stores only a portion (e.g., a transient portion) of the audio stream received within a predetermined duration (e.g., 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, or 2 seconds) before the current time, and stores only a portion (e.g., a transient portion) of the video stream received within a predetermined duration before the current time. For example, memory buffer 804 is a first - in - first - out (FIFO) buffer (e.g., a circular buffer) (e.g., optionally cropped to include the user's head) that includes the last x seconds of the audio stream and the last x seconds of the video stream, where x is the predetermined duration before the current time.
[0221] System 800 includes an audiovisual calling module 806. The audiovisual calling module 806 is configured to determine whether a user's visual attention is directed towards an electronic device while the user is speaking, based on an audio stream and a video stream. In some embodiments, the audiovisual calling module 806 makes the determination by processing the stored content in the memory buffer 804 based on, for example, a portion of the audio stream and the video stream each received within a predetermined duration prior to the current time.
[0222] In some embodiments, determining whether a user's visual attention is directed towards an electronic device while the user is speaking includes determining whether the user's visual attention is directed towards a display of the electronic device (e.g., the display of device 900 in FIGS. 9 and 10A - 10B). In some embodiments, determining whether a user's visual attention is directed towards an electronic device while the user is speaking includes determining whether the user's visual attention is directed towards an affordance (e.g., an affordance representing a digital assistant) presented by the electronic device. In some embodiments, determining that a user's visual attention is directed towards an electronic device while the user is speaking includes determining that the user's line of sight is directed towards the electronic device while the user is speaking. In some embodiments, determining that a user's visual attention is directed towards an electronic device while the user is speaking includes determining that the user (e.g., the user's posture (e.g., head posture)) is facing the electronic device while the user is speaking.
[0223] In some embodiments, the audiovisual calling module 806 implements a machine learning model trained to determine whether the user's visual attention is directed towards the electronic device while the user is speaking. The machine learning model is implemented, for example, as a neural network (e.g., a recurrent neural network, a convolutional neural network, a feedforward neural network, etc.) and is configured to receive as input the contents of the memory buffer 804 (e.g., a portion of the audio stream and the video stream). The neural network is trained using a labeled training data set that includes different audiovisual training data, e.g., audio data and video data received simultaneously. Each element of the audiovisual training data is labeled as follows. ● While the user is speaking, the user's visual attention is directed towards the electronic device (speaking while looking). ● While the user is speaking, the user's visual attention is not directed towards the electronic device (speaking without looking). ● While the user is not speaking, the user's visual attention is directed towards the electronic device (looking but not speaking). ● While the user is not speaking, the user's visual attention is not directed towards the electronic device (not looking and not speaking), or ● The user is not visible in the audiovisual data (regardless of whether the audio stream includes user speech). In this way, the audiovisual calling module 806 is trained to classify the audio stream and the video stream as representing (1) speaking while looking, (2) speaking without looking, (3) looking but not speaking, (4) not looking and not speaking, or (5) the user not being visible.
[0224] Furthermore, such training can enable the audiovisual invocation module 806 to implicitly learn features in the audio stream and video stream indicating that the user is looking at and speaking to the electronic device, e.g., a correlation between the user's mouth movements and the utterance input in the audio stream. For example, the audiovisual invocation module 806 can use parameters (e.g., learned weights) of a machine learning model representing the correlation between the user's mouth movements and the utterance input to process representations of a portion of the audio stream and a portion of the video stream (each received within a predetermined duration prior to the current time).
[0225] The following describes the technique used by the audiovisual invocation module 806 to determine whether the user's visual attention is directed towards the electronic device while the user is speaking (referred to briefly as "whether the user is looking and speaking").
[0226] FIG. 8B shows video frames stored in the memory buffer 804 at the current time according to various examples. As shown, a portion of the video stream received within a predetermined duration prior to the current time includes a plurality of video frames 852, 854, 856, 858, 860, 862, and 864. This example shows that the memory buffer 804 includes seven video frames, but the number of video frames stored in the memory buffer 804 can vary based on a predetermined duration (e.g., 1 second or 1.5 seconds), the camera frame rate, and / or the amount of memory allocated to the memory buffer 804. The video frames 852, 854, 856, 858, 860, 862, and 864 respectively correspond to times T1, T2, T3, T4, T5, T6, and T7, e.g., the times at which the video frames were captured.
[0227] In some embodiments, determining whether a user is looking and speaking involves determining an individual reliability score for each of video frames 852-864 to obtain a plurality of respective reliability scores S1-S7. Each individual reliability score indicates whether the user is looking and speaking for the individual video frame. For example, if the reliability score exceeds a threshold, the reliability score indicates that the user is looking and speaking. If the reliability score does not exceed the threshold, the reliability score indicates that the user is speaking without looking (or looking without speaking, or not looking and not speaking, or the user is not visible). As another example, the value of the reliability score indicates the reliability level (e.g., high, medium, or low) that the user is looking and speaking.
[0228] In some embodiments, for each of video frames 852, 854, 856, 858, 860, 862, and 864, the audiovisual call module 806 determines other reliability scores indicating (1) whether the user is not looking and speaking, (2) whether the user is looking and not speaking, (3) whether the user is not looking and not speaking, and (4) whether the user is not visible, respectively. In some embodiments, the audiovisual call module 806 normalizes the reliability scores of each video frame so that the sum is 1.
[0229] In some embodiments, determining a reliability score (e.g., S4) for a video frame (e.g., video frame 858) is based on processing the video frame, for example, using a machine learning model, to determine an initial reliability score of the video frame (e.g.,
Number
Number
Number
[0230] In some embodiments, determining the reliability score for a video frame further includes adjusting the initial reliability score to determine the reliability score based on processing one or more subsequent video frames (e.g., corresponding to the time(s) after T4). For example, at time T7 in FIG. 8B, the machine learning model processes the subsequent video frames 860, 862, and 864 and also uses the context information obtained by processing the audio content stored in the memory buffer 804 at time T7 to determine the initial reliability score of the video frame 858
Number
[0231] As an example of adjusting the initial reliability score of video frame 858,
Number
Number
Number
Number
[0232] In some embodiments, the audiovisual call module 806 determines that the user is looking and speaking by determining that the reliability score for any video frame in the memory buffer 804 exceeds a threshold, or by determining that the average reliability score for the video frames in the memory buffer 804 exceeds a threshold
[0233] In some embodiments, the audiovisual call module 806 determines that the user is looking and speaking by determining that the reliability score of a particular video frame exceeds a threshold. In some embodiments, a particular video frame occupies a preselected position within the memory buffer 804, such as the earliest position (occupied by video frame 852), the middle position (occupied by video frame 858), or the latest position (occupied by video frame 864). Determining whether the reliability score for the latest video frame 864 exceeds the threshold can quickly determine whether the user is looking and speaking, but the determination may not be very accurate, for example, because the determination depends on the initial reliability score of the video frame. In contrast, determining whether the reliability score for the earliest video frame 852 exceeds the threshold can more accurately determine whether the user is looking and speaking (e.g., because the reliability score is based on processing subsequent video frames), but the determination can be relatively slow. Thus, the preselected video frame position for determining whether the user is looking and speaking can vary to balance the appropriate accuracy and speed considerations for a particular implementation of the system 800. In some embodiments, the machine learning model is configured to optimize (e.g., minimize) a loss function associated with the preselected video frame position, such as minimizing the loss function associated with the middle video frame 858.
[0234] System 800 includes a post - processing module 808. In accordance with a determination that a user is looking and speaking, the post - processing module 808 is configured to identify a portion of an audio stream that includes a user utterance (an utterance intended by the device) addressed to the electronic device. In some embodiments, identifying a portion of the audio stream to include an utterance intended by the device includes determining that a time (e.g., T4) corresponding to a video frame (e.g., occupying a pre - selected position) is the start time of a portion of the audio stream in accordance with a determination that a reliability score for the video frame exceeds a threshold. In some embodiments, determining the start time further includes determining that a reliability score for a previous (e.g., immediately preceding) video frame does not exceed the threshold. Thus, the post - processing module 808 can determine the start time by determining that a reliability score corresponding to a pre - selected video frame position exceeds the threshold.
[0235] In some embodiments, identifying a portion of the audio stream to include an utterance intended by the device includes determining that a reliability score for a first video frame exceeds a threshold and determining that a reliability score for a second video frame (e.g., occupying a pre - selected position) consecutive to the first video frame is below the threshold. In accordance with such determinations, the post - processing module 808 determines that a time corresponding to the second video frame is the end time of a portion of the audio stream. Thus, the post - processing module 808 can determine the end time by determining that a reliability score corresponding to a pre - selected video frame position is below the threshold. In some embodiments, the post - processing module 808 implements, additionally or alternatively, an utterance - end - specifying technique known in the art to determine the end time. The reliability scores described with respect to identifying the start and end times can be an initial reliability score for the video frame, an adjusted reliability score for the video frame, or a final reliability score for the video frame (described in detail below).
[0236] Those skilled in the art will understand that the post - processing module 808 can implement various other techniques for using a reliability score (or scores) (e.g., an initial reliability score (or scores), an adjusted reliability score (or scores), a final reliability score (or scores)) to determine a start time and / or an end time. As an example, in accordance with determining that each video frame in the memory buffer 804 has an individual reliability score that exceeds a threshold, the post - processing module 808 determines the start time as the time corresponding to the earliest video frame (or an intermediate video frame, or the latest video frame) in the memory buffer 804. Similarly, in accordance with determining that each video frame in the memory buffer 804 has an individual reliability score that is below the threshold, the post - processing module 808 determines the end time as the time corresponding to the earliest video frame (or an intermediate video frame, or the latest video frame) in the memory buffer 804. As another example, in accordance with determining that the respective reliability scores for a consecutive number (e.g., two, three, or four) of video frames within the memory buffer 804 each exceed the threshold, the post - processing module 808 determines the start time as the time corresponding to the earliest video frame among the consecutive video frames. Similarly, in accordance with determining that the respective reliability scores for a consecutive number (e.g., two, three, or four) of video frames within the memory buffer 804 each are below the threshold, the post - processing module 808 determines the end time as the time corresponding to the earliest video frame among the consecutive video frames. As another example, in accordance with determining that the average reliability score for a consecutive number (e.g., two, three, or four) of video frames within the memory buffer 804 exceeds the threshold, the post - processing module 808 determines the start time as the time corresponding to the earliest video frame among the consecutive video frames. Similarly, in accordance with determining that the average reliability score for a consecutive number (e.g., two, three, or four) of video frames within the memory buffer 804 is below the threshold, the post - processing module 808 determines the end time as the time corresponding to the earliest video frame among the consecutive video frames.Those skilled in the art will understand that other variations are possible for determining the start time and / or end time using the reliability score(s), and that these fall within the scope of the present disclosure.
[0237] The post-processing module 808 is further configured to cause a digital assistant operating on the device 900 to initiate a task based on an identified portion of the audio stream. For example, the post-processing module 808 causes the digital assistant to process an identified portion of the audio stream as described with respect to FIGS. 7A-7C. In some embodiments, the device 900 provides an output indicating the initiated task.
[0238] FIG. 9 shows a device 900 that provides an output according to a determination that the visual attention of user 904 is directed towards the device 900 while user 904 is speaking, according to various examples. The device 900 is implemented as device 104, 122, 200, 400, or 600 and includes one or more cameras 902. The device 900 further at least partially implements the digital assistant system 700 as described with respect to FIGS. 7A-7C. FIGS. 9 and 10A-10B show the device 900 being implemented as a particular type of device, but the device 900 may be implemented as another type of device, such as a smartphone, laptop computer, desktop computer, smartwatch, television, smart speaker, head-mounted device, smart home appliance, tablet device, or vehicle head unit.
[0239] In FIG. 9, according to the techniques discussed above, the system 800 identifies the utterance input "read message" of user 904 as the intended utterance of the device. Accordingly, the digital assistant initiates the task of reading user 904's message and provides an output (e.g., in an audio and / or displayed format) "Ok, your first message from Max is 'hello'."
[0240] In this way, device 900 identifies a portion of the audio stream to include the speech intended for the device without detecting a speech trigger (e.g., "Hey Siri", "Siri", "Hey Assistant", "Ok, computer", etc.) to initiate a digital assistant session and without relying on other explicit instructions (e.g., user selection of a hardware button, user selection of a displayed affordance) addressed to device 900. This can provide a more natural and efficient interaction with device 900.
[0241] In some embodiments, system 800 includes an identification module 810. In some embodiments, identification module 810 is configured to receive an audio stream and identify user 904 based on the audio stream, e.g., according to speech recognition techniques known in the art. In some embodiments, identification module 810 is configured to receive a video stream and identify user 904 based on the video stream, e.g., according to image / face recognition techniques known in the art. In some embodiments, the output (e.g., "Ok, your first message from Max is "hello".") is based on the identified user. For example, the output is personalized for the identified user by reading the identified user's messages.
[0242] In some embodiments, the identification module 810 identifies the user 904 according to a determination by the digital assistant that fulfilling the user's intent includes retrieving or modifying the personal data of the user 904. Exemplary personal data includes the messages, emails, photos, videos, calendar information, memos, financial information, health information, home security information (e.g., whether the device is locked, whether the appliances are on, the settings of the appliances), voice recordings, and any other potentially confidential information that the user 904 may not wish to be made public to others. In some embodiments, the identification module 810 combines the confidence level associated with the voice identification of the user 904 and the confidence level associated with the face recognition of the user 904 to identify the user 904. For example, if the reliability scores associated with voice identification and face identification are each insufficient on their own to confidently identify the user 904, e.g., if voice identification and face identification each show a moderate confidence level in the same user 904, the combination of the scores may be able to identify the user 904 more confidently. In some embodiments, in accordance with identifying the user 904, the identification module 810 causes the digital assistant to provide an output personalized for the identified user 904.
[0243] In some embodiments, in accordance with a determination that the user's visual attention is not directed towards the electronic device while the user is speaking (the user is not looking and is speaking), the post-processing module 808 stops identifying a portion of the audio stream to include the utterance intended by the device. For example, the audiovisual call module 806 determines that the user is not looking and is speaking by determining that a reliability score (e.g., for a video frame occupying a preselected position) indicating whether the user is not looking and is speaking exceeds a threshold.
[0244] In some embodiments, in accordance with a determination that the user's visual attention is directed to the electronic device while the user is not speaking (the user is looking but not speaking), the post-processing module 808 ceases to identify a portion of the audio stream to include the utterance intended by the device. For example, the audiovisual call module 806 determines that the user is looking and not speaking by determining that a reliability score (e.g., for a video frame occupying a preselected position) indicating whether the user is looking and not speaking exceeds a threshold.
[0245] In some embodiments, in accordance with a determination that the user's visual attention is not directed to the electronic device while the user is not speaking (the user is not looking and not speaking), the post-processing module 808 ceases to identify a portion of the audio stream to include the utterance intended by the device. For example, the audiovisual call module 806 determines that the user is not looking and not speaking by determining that a reliability score (e.g., for a video frame occupying a preselected position) indicating whether the user is not looking and not speaking exceeds a threshold.
[0246] In some embodiments, in accordance with a determination that the user is not visible in a first portion of the video stream, the post-processing module 808 ceases to identify a portion of the audio stream to include the utterance intended by the device. For example, the audiovisual call module 806 determines that the user is not visible by determining that a reliability score (e.g., for a video frame occupying a preselected position) indicating whether the user is not visible exceeds a threshold.
[0247] FIGS. 10A-10B illustrate determining whether the user's visual attention is directed to device 900 while the same user is speaking in a multi-user environment, according to various examples.
[0248] In FIGS. 10A-10B, device 900 (e.g., system 800) receives an audio stream and a video stream simultaneously, as described above with respect to FIG. 8A. In some embodiments, the audio stream includes the speech of first user 1000 and the speech of second user 1002. In some embodiments, the video stream includes the video of first user 1000 and the video of second user 1002. For example, in a multi-user environment, device 900 can detect multiple users (e.g., via one or more cameras) and detect the speech of multiple users (e.g., via one or more microphones). As described below, device 900 can determine whether the visual attention of a user is directed towards device 900 while the same user is speaking, in order to provide a response related to the user and avoid accidentally responding to speech not addressed to device 900.
[0249] In some embodiments, device 900 identifies a first portion (e.g., a cropped portion) of the video stream that includes first user 1000 and a second portion (e.g., a cropped portion) of the video stream that includes second user 1002, as described above with respect to preprocessing module 802.
[0250] Based on the first portions of the audio stream and the video stream, device 900 determines whether the first user 1000 is looking and talking, for example, according to the techniques described above with respect to FIGS. 8A-8B. Similarly, based on the second portions of the audio stream and the video stream, device 900 determines whether the second user 1002 is looking and talking, for example, according to the techniques described above with respect to FIGS. 8A-8B. For example, device 900 stores a portion of the audio stream received within a predetermined duration prior to the current time and a first portion of the video stream received within a predetermined duration prior to the current time (e.g., the last x seconds of the cropped video stream) in a first instance of memory buffer 804. Device 900 further stores a portion of the audio stream and a second portion of the video stream received within a predetermined duration prior to the current time in a second instance of memory buffer 804. Next, the audiovisual call module 806 processes the stored content of each instance of memory buffer 804 separately as described above, thereby determining a reliability score (s) for the video frame (s) indicating whether the user is looking and talking for each user. In some embodiments, processing the stored content of each instance of memory buffer 804 further includes determining other reliability scores for the video frame (s) indicating, for each user, (1) whether the user is not looking and talking, (2) whether the user is looking and not talking, (3) whether the user is not looking and not talking, and (4) whether the user is not visible.
[0251] Based on the first part of the audio stream and the video stream, according to the determination that the first user 1000 is watching and speaking, the device 900 identifies the first part of the audio stream that includes the speech of the first user 1000 addressed to the device 900. Based on the second part of the audio stream and the video stream, according to the determination that the second user 1002 is watching and speaking, the device 900 identifies the second part of the audio stream that includes the speech of the second user 1002 addressed to the device 900. For example, according to the technology described above, the post-processing module 808 identifies the start time and / or end time of the first part of the audio stream (and / or the second part of the audio stream).
[0252] In FIG. 10A, the device 900 further provides a first output based on the processing of the first part of the audio stream identified as including the speech intended by the device from the user 1000. In some embodiments, providing the first output includes using a digital assistant to process the first part of the audio stream. For example, since it is determined that the first user 1000 is looking at and speaking to the device 900, the digital assistant processes the speech input "Read message" of the first user 1000 and provides the output "Ok, your first message from Max is 'hello'." In the example of FIG. 10A, the second user 1002 is determined to be looking at the device 900 and not speaking. Therefore, the device 900 stops identifying any second part of the audio stream and does not provide any output to the second user 1002.
[0253] In an example where the device 900 identifies the second part of the audio stream such that it includes the speech of the second user 1002 addressed to the device 900, the device 900 similarly provides a second output based on the processing of the second part of the audio stream. In some embodiments, providing the second output includes using a digital assistant to process the second part of the audio stream.
[0254] In some embodiments, the identification module 810 identifies a first user 1000 based on a first portion (e.g., a cropped portion) of the audio stream and / or the video stream. In some embodiments, the first output is based on the identified first user 1000. For example, in FIG. 10A, the first output "Ok, your first message from Max is 'hello'." is personalized for the identified first user 1000. In some embodiments, the identification module 810 identifies a second user 1002 based on a second portion (e.g., a cropped portion) of the audio stream and / or the video stream. In some embodiments, the second output is based on the identified second user 1002 and is personalized, for example, for the identified second user 1002.
[0255] In some embodiments, based on a first portion of the audio stream and the video stream, in accordance with the determination that the first user 1000 is viewing but not speaking, the device 900 stops identifying the first portion of the audio stream that includes the utterances of the first user 1000 addressed to the device 900. Similarly, in some embodiments, based on a second portion of the audio stream and the video stream, in accordance with the determination that the second user 1002 is viewing but not speaking, the device 900 stops identifying the second portion of the audio stream that includes the utterances of the second user 1002 addressed to the device 900. Similarly, the device 900 stops identifying the first (or second) portion of the audio stream in accordance with the determination that the first user 1000 (or the second user 1002) (1) is not viewing and is speaking, (2) is not viewing and is not speaking, or (3) is not visible, based on the first (or second) portion of the audio stream and the video stream.
[0256] Referring to FIG. 10B, the device 900 determines that the visual attention of the first user 1000 is directed to the device 900 while the second user 1002 is speaking. For example, the post-processing module 808 determines that the first video frame indicates that the first user 1000 is looking and not speaking, and the second video frame indicates that the second user 1002 is not looking and is speaking, and the first and second video frames correspond to the same time. In accordance with such a determination, the device 900 stops identifying any part of the audio stream that contains the utterance intended by the device, and thus stops providing any response to the utterance input. For example, the device 900 does not misinterpret the utterance "What's for dinner?" directed at the device 900 by the second user 1002 as being for the device 900, even though the visual attention of the first user 1000 is directed to the device 900 while the device 900 is receiving the utterance input. Similarly, in accordance with the determination that the visual attention of the second user 1002 is directed to the device 900 while the first user 1000 is speaking (e.g., the second user 1002 is looking but not speaking while the first user 1000 is speaking but not looking), the device 900 stops identifying any part of the audio stream that contains the utterance intended by the device. In this way, to provide a response to the utterance input, the device 900 determines whether the same user is looking and speaking. This can prevent the device 900 from providing an incorrect response, for example, from incorrectly outputting search results for a restaurant in response to "What's for dinner?" in FIG. 10B.
[0257] In some embodiments, the reliability score indicating whether the user is looking and talking (e.g., for a video frame) is a first type of reliability score (a visual score). In some embodiments, after determining the visual score, the post - processing module 808 determines a final reliability score indicating whether the user is looking and talking based on the visual score and one or more other types of reliability scores indicating whether the user's visual attention is directed towards the device 900. In some embodiments, the final reliability score is for the video frame of the same video stream as the visual score. For example, referring to FIG. 8B, the post - processing module 808 determines the final reliability score for the video frame 858 based on the visual score S4 and one or more other types of reliability scores.
[0258] Examples of other types of reliability scores are discussed here. Determining the final reliability score based on other types of reliability scores can potentially refine the visual score based on potentially relevant factors such as the user's line of sight, the user's posture, the relative motion between the user and the device 900, the user's gestures, etc., as described below, to more accurately determine whether the user is looking and talking.
[0259] In some embodiments, the second type of reliability score (gaze score) indicates whether the user's gaze is directed towards device 900. For example, a gaze model (e.g., a machine learning model) implemented in post - processing module 808 or on a device external to device 900 processes the video stream to determine the gaze score. In some embodiments, the gaze model is similar to the machine learning model described above with respect to the audiovisual call module 806. For example, the gaze model determines the gaze score for a video frame (or for a particular time). In some embodiments, having a gaze score above a threshold indicates that the user's gaze is directed towards device 900, while having a gaze score below the threshold indicates that the user's gaze is not directed towards device 900.
[0260] In some embodiments, if the gaze score indicates that the user's gaze is directed towards device 900, post - processing module 808 increases the final reliability score with respect to the audiovisual score. Similarly, in some embodiments, if the gaze score indicates that the user's gaze is not directed towards device 900, post - processing module 808 decreases (or does not change) the final reliability score with respect to the audiovisual score. In some embodiments, the gaze score corresponds to the same time as the audiovisual score. For example, if the gaze score for time T4 in FIG. 8B indicates that the user's gaze is directed towards device 900, post - processing module 808 increases the final reliability score for video frame 858.
[0261] In some embodiments, a third type of reliability score (posture score) indicates whether the user's posture (e.g., head posture) is facing the device 900. For example, a posture model (e.g., a machine learning model) implemented in the post-processing module 808 or on a device external to the device 900 processes the video stream to determine the posture score. In some embodiments, the posture model is similar to the machine learning model described above with respect to the audiovisual call module 806. For example, the posture model determines a posture score for a video frame (or for a particular time). In some embodiments, having a posture score above a threshold indicates that the user's posture is facing (or starting to face) the device 900, while having a posture score below the threshold indicates that the user's posture is not facing (or starting to face) the device 900.
[0262] In some embodiments, the post-processing module 808 increases the final reliability score for the audiovisual score if the posture score indicates that the user's posture is facing (or starting to face) the device 900. Similarly, in some embodiments, the post-processing module 808 decreases (or does not change) the final reliability score for the audiovisual reliability score if the posture score indicates that the user's posture is not facing (or starting to face) the device 900. In some embodiments, the posture score corresponds to the same time as the audiovisual score. For example, if the posture score at time T4 in FIG. 8B indicates that the user's posture is not facing the device 900, the post-processing module 808 decreases the final reliability score for the video frame 858.
[0263] In some embodiments, a fourth type of reliability score (gesture score) indicates whether a gesture of a predetermined type of user is detected. Exemplary gestures of a predetermined type of user include types of hand gestures (e.g., waving gesture, pointing gesture, raising hand gesture, grasping gesture, pushing gesture, etc.) and types of finger gestures (e.g., movement of a finger (singular or plural) in a particular manner). For example, a gesture model (e.g., a machine learning model) implemented in the post - processing module 808 or on a device external to the device 900 processes the video stream to determine the gesture score. In some embodiments, the gesture model is similar to the machine learning model described above with respect to the audiovisual call - out module 806. For example, the gesture model determines the gesture score for a video frame (or for a particular time). In some embodiments, having a gesture score above a threshold indicates that a gesture of a predetermined type is detected, while having a gesture score below the threshold indicates that a gesture of a predetermined type is not detected.
[0264] In some embodiments, if the gesture score indicates that a gesture of a predetermined type is detected, the post - processing module 808 increases the final reliability score with respect to the audiovisual reliability score. Similarly, in some embodiments, if the gesture score indicates that a gesture of a predetermined type is not detected, the post - processing module 808 decreases (or does not change) the final reliability score with respect to the audiovisual score. In some embodiments, the gesture score corresponds to the same time as the audiovisual score. For example, if the gesture score for time T4 in FIG. 8B indicates that a gesture of a predetermined type of user is detected (e.g., the user is performing a gesture at time T4), the post - processing module 808 increases the final reliability score for the video frame 858.
[0265] In some embodiments, a fifth type of reliability score (relative motion score) indicates the relative motion between the user and device 900, e.g., whether the user is moving towards device 900. For example, a relative motion model (e.g., a machine learning model) implemented in post-processing module 808 or on a device external to device 900 processes the video stream and / or audio stream to determine the relative motion score. In some embodiments, the relative motion model is similar to the machine learning model described above with respect to the audiovisual call module 806. For example, the relative motion model determines, for a video frame (or for a particular time), a relative motion score indicating whether the user is moving towards device 900. In some embodiments, having a relative motion score above a threshold indicates that the user is moving towards device 900, while having a relative motion score below the threshold indicates that the user is not moving towards device 900.
[0266] In some embodiments, if the relative motion score indicates that the user is moving towards device 900, post-processing module 808 increases the final reliability score with respect to the audiovisual score. Similarly, in some embodiments, post-processing module 808 decreases (or does not change) the final reliability score with respect to the audiovisual score if the relative motion score indicates that the user is not moving towards device 900. In some embodiments, the relative motion score corresponds to the same time as the audiovisual score. For example, if the relative motion score for time T4 in FIG. 8B indicates that the user is moving towards device 900 (e.g., the user is moving towards device 900 at time T4), post-processing module 808 increases the final reliability score for video frame 858.
[0267] In some embodiments, determining the final reliability score includes determining whether the content of the audio stream (e.g., content by words) corresponds to a digital assistant request, e.g., whether the audio stream includes requests typically spoken to a digital assistant. In some embodiments, the post-processing module 808 determines whether the content corresponds to a digital assistant request based on determining the type of the content, e.g., by using the natural language processing capabilities of the digital assistant. For example, if the content is of a first type (e.g., a command or a question), the post-processing module 808 determines that the content corresponds to a digital assistant request and / or increases the probability that the content corresponds to a digital assistant request. If the content is of a second type (e.g., an opinion), the post-processing module 808 determines that the content does not correspond to a digital assistant request and / or decreases the probability that the content corresponds to a digital assistant request. In some embodiments, the post-processing module 808 determines whether the content corresponds to a digital assistant request based on determining whether the content corresponds to the vocabulary associated with the domain of the digital assistant, e.g., by using the natural language processing capabilities of the digital assistant. For example, if the content includes at least a portion (e.g., at least a threshold number) of the words in the vocabulary (e.g., the words "food", "drink", "dish", "hungry", "eat", "pizza", "fast food", and "meal" in the restaurant reservation domain 762), the post-processing module 808 determines that the content corresponds to a digital assistant request and / or increases the probability that the content corresponds to a digital assistant request. If the content does not include at least a portion of the words in the vocabulary (e.g., includes less than a threshold number), the post-processing module 808 determines that the content does not correspond to a digital assistant request and / or decreases the probability that the content corresponds to a digital assistant request.
[0268] In some embodiments, the post - processing module 808 increases the final reliability score for the audiovisual score according to the determination that the content corresponds to a digital assistant request. In some embodiments, the post - processing module 808 decreases (or does not change) the final reliability score for the audiovisual score according to the determination that the content does not correspond to a digital assistant request. In some embodiments, the content corresponds to the same time as the audiovisual score. For example, the device 900 receives the content within a predetermined time window (e.g., 1, 2, 3, 4, or 5 seconds) near the time corresponding to the audiovisual score. For example, if the post - processing module 808 determines that the device 900 has received content corresponding to a digital assistant request near time T4 in FIG. 8B, the post - processing module 808 increases the final reliability score of the video frame 858. In this way, although the audiovisual score indicates an insufficient confidence level that the user is watching and speaking (e.g., the device 900 is only partially watching the user while the user is speaking), if the content of the user's speech corresponds to a typical digital assistant request, the device 900 can still accurately identify and respond to the intended speech of the user's device.
[0269] In some embodiments, determining the final reliability score includes determining, by device 900, the actions currently being performed by device 900, e.g., when determining the final reliability score. In some embodiments, in accordance with a determination that the action is of a first type, post-processing module 808 increases the final reliability score. In some embodiments, in accordance with a determination that the action is of a second type different from the first type, post-processing module 808 decreases the final reliability score. An action of the first type can indicate an increase in the likelihood of user interaction with the digital assistant. Exemplary actions of the first type include displaying a notification (e.g., a new message notification, a system notification, a notification that a timer has expired, a phone call incoming notification), actively executing the digital assistant (e.g., displaying the digital assistant user interface), and displaying a predetermined type of user interface (e.g., a home screen user interface, a user interface of a predetermined application). An action of the second type can indicate a decrease in the likelihood of user interaction with the digital assistant. Exemplary actions of the second type include outputting media content (e.g., a movie, a song, an audio book) and actively executing a communication application (e.g., a call application, a video call application).
[0270] In some embodiments, the post - processing module 808 determines a final reliability score based on data representing user interactions with the digital assistant. In some embodiments, the data indicates a high likelihood of user interaction with the digital assistant, such as, for example, the device 900 actively running the digital assistant (e.g., displaying the digital assistant user interface), the digital assistant frequently starting on the device 900 (e.g., starting more than a threshold number of times within a predetermined duration), the digital assistant having recently (e.g., within a predetermined duration before the current time) output a response to the user, and / or the digital assistant having recently ended (e.g., the device 900 having recently stopped displaying the digital assistant user interface). In some embodiments, the data indicates a decrease in the likelihood of user interaction with the digital assistant, such as, for example, the device 900 not currently displaying the digital assistant user interface, the digital assistant not frequently starting on the device 900 (e.g., starting less than a threshold number of times within a predetermined duration), the digital assistant not having provided an output within a predetermined duration before the current time (e.g., within the past 1 hour, within the past 1 day), and / or the digital assistant having been last started more than a predetermined duration before the current time (e.g., more than 1 month ago). In some embodiments, in accordance with a determination that the data indicates an increase in the likelihood of user interaction with the digital assistant, the post - processing module 808 increases the final reliability score. In some embodiments, in accordance with a determination that the data indicates a decrease in the likelihood of user interaction with the digital assistant, the post - processing module 808 decreases the final reliability score.
[0271] In some embodiments, the post - processing module 808 determines whether the final reliability score exceeds a threshold, e.g., whether the user is looking and speaking. For example, the post - processing module 808 determines whether the final reliability score for a video frame occupying a pre - selected position (e.g., the earliest position, the middle position, or the latest position) in the memory buffer 804 exceeds the threshold.
[0272] In some embodiments, the post - processing module 808 adjusts a threshold and compares the final confidence score to the adjusted threshold. As will be described below, adjusting the threshold can make it more or less likely that the system 800 determines whether the user is looking and speaking, which can then increase the system accuracy.
[0273] In some embodiments, in accordance with the determination that the action currently being performed by the device 900 is of a first type, the post - processing module 808 decreases the threshold, and in accordance with the determination that the action is of a second type, the post - processing module 808 increases the threshold. As another example, the post - processing module 808 adjusts the threshold based on data representing a user interaction with a digital assistant (e.g., decreases the threshold if the data corresponds to an increase in the final confidence score and increases the threshold if the data corresponds to a decrease in the final confidence score). As another example, the post - processing module 808 adjusts the threshold based on at least a portion of other types of confidence scores (e.g., a gaze score, a pose score, a gesture score, a relative motion score). For example, if the other types of confidence scores correspond to an increase (or decrease) in the final confidence score, the post - processing module 808 decreases (or increases) the threshold, respectively.
[0274] In some embodiments, the post - processing module 808 reduces the threshold in accordance with a determination that the user is alone within the physical environment associated with the device 900. The physical environment can be, for example, the room in which the device 900 is located, a predetermined volume around the device 900 (e.g., within 10, 15, 20, or 25 feet of the device 900), or the interior of a vehicle (e.g., if the device 900 is within a vehicle). In some embodiments, the post - processing module 808 determines that the user is alone within the physical environment based on analyzing an audio stream (e.g., by detecting sounds (e.g., footsteps) corresponding to a single user), based on detecting user device(s) (e.g., detecting only user device(s) belonging to a single user within the physical environment), based on analyzing a video stream (e.g., determining that the video stream includes a single user), or based on a combination or sub - combination thereof. In this way, the system 800 can be more likely to interpret utterances spoken when the user is alone as being addressed to the device 900.
[0275] In some embodiments, in accordance with a determination that the final confidence score exceeds a (optionally adjusted) threshold (e.g., the user is looking and speaking), the post - processing module 808 identifies a portion of the audio stream that includes the utterance intended by the device. For example, instead of using visual score(s) to determine start time and / or end time, the post - processing module 808 uses the final confidence score(s) in a similar way to determine start time and / or end time.
[0276] In some embodiments, identifying a portion of an audio stream to include the utterance intended by the device includes determining an initial start time of a portion of the audio stream to be a time corresponding to a visual score or a final confidence score. For example, the post-processing module 808 determines the initial start time to be a time (e.g., T4) corresponding to a video frame (e.g., video frame 858) having a visual score and / or a final confidence score that is determined to exceed a (optionally adjusted) threshold. In some embodiments, the post-processing module 808 adjusts the initial start time based on a third type of confidence score (posture score). For example, assume that the post-processing module 808 determines that the posture score has a relatively high value at time T1 in FIG. 8B, and for example, indicates that the user has started to turn towards the device at time T1. Further, if the audiovisual call module 806 determines that the user is speaking at time T1 (e.g., if a confidence score indicating whether the user is not looking and is speaking at time T1 exceeds a threshold), the post-processing module 808 adjusts the start time to T1. In this way, if the user starts to turn towards the device 900 (but has not yet faced the device 900) and starts to provide the utterance intended by the device, the system 800 can identify the correct start time of the utterance intended by the device. For example, relying on the visual score (rather than the posture score) to determine the start time may inadvertently result in determining that the start time is time T4, for example, in the middle of the utterance intended by the user for the device.
[0277] In some embodiments, the device 900 provides an output based on processing a portion of the audio stream identified to include the utterance intended by the device, as shown, for example, in FIG. 9.
[0278] In some embodiments, in accordance with a determination that the final confidence score does not exceed a threshold, the post-processing module 808 stops identifying portions of the audio stream to include the utterance intended by the device.
[0279] In the above, the determination of the final reliability score for one user and the adjustment of the threshold have been described. However, it will be understood that such techniques are equally applicable to determining the final reliability score for multiple users and adjusting the threshold. For example, the audiovisual call module 806 processes the stored content of different instances of the memory buffer 804 (each corresponding to a different user) separately to determine the audiovisual score for different users. Then, for each user, the post-processing module 808 separately determines the final reliability score, and optionally, separately adjusts the individual threshold, and separately compares the final reliability score with the (optionally adjusted) individual threshold to determine whether each of the different users is looking and talking. In accordance with the determination that the user is looking and talking, the post-processing module 808 identifies a portion of the audio stream that includes the utterance intended by the device. Then, as described, the device 900 provides an output based on processing the identified portion of the audio stream. 5. A process for determining whether a user's visual attention is directed towards an electronic device while the user is talking
[0280] FIG. 11 shows a process 1100 for determining whether a user's visual attention is directed to an electronic device while the user is speaking, according to various examples. The process 1100 is executed using, for example, one or more electronic devices implementing a digital assistant. In some embodiments, the process 1100 is executed using a client-server system (e.g., system 100), and the blocks of the process 1100 are divided in any manner between a server (e.g., DA server 106) and a client device (e.g., device 900). In other embodiments, the blocks of the process 1100 are divided between a server and a plurality of client devices (e.g., a mobile phone and a smartwatch). Thus, although some parts of the process 1100 are described herein as being executed by a particular device of a client-server system, it will be understood that the process 1100 is not so limited. In other embodiments, the process 1100 is executed using only a client device (e.g., user device 104, or 900) or only a plurality of client devices. In the process 1100, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in combination with the process 1100.
[0281] In block 1102, an audio stream and a video stream are received simultaneously, for example, by device 900.
[0282] In block 1104, based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within a predetermined duration before the current time, it is determined (e.g., by the audiovisual call module 806) whether the user's visual attention is directed to the electronic device (e.g., device 900) while the user (e.g., user 904, 1000, or 1002) is speaking.
[0283] In some embodiments, determining whether the user's visual attention is directed to the electronic device while the user is speaking includes determining whether the user's visual attention is directed to the display of the electronic device. In some embodiments, determining whether the user's visual attention is directed to the electronic device while the user is speaking includes determining whether the user's visual attention is directed to an affordance displayed by the electronic device.
[0284] In some embodiments, determining that the user's visual attention is directed to the electronic device while the user is speaking includes determining that the user's line of sight is directed to the electronic device while the user is speaking. In some embodiments, determining that the user's visual attention is directed to the electronic device while the user is speaking includes determining that the user is facing the electronic device while the user is speaking.
[0285] In some embodiments, the first portion of the audio stream and the first portion of the video stream are stored in a memory buffer (e.g., memory buffer 804), and determining whether the user's visual attention is directed to the electronic device while the user is speaking includes processing the stored content of the memory buffer.
[0286] In some embodiments, the first portion of the video stream includes a plurality of video frames (e.g., video frames 852, 854, 856, 858, 860, 862, and 864). In some embodiments, determining whether the user's visual attention is directed to the electronic device while the user is speaking includes determining an individual reliability score (e.g., reliability scores S1, S2, S3, S4, S5, S6, and S7) for each video frame of the plurality of video frames to obtain a plurality of respective reliability scores, where the individual reliability score indicates whether the user's visual attention is directed to the electronic device while the user is speaking for the individual video frame.
[0287] In some embodiments, each of the plurality of reliability scores includes a first reliability score for a first video frame among the plurality of video frames (e.g., reliability score S4 for video frame 858), and the first video frame corresponds to a first time (e.g., time T4). In some embodiments, determining the first reliability score includes determining an initial first reliability score (e.g., initial reliability score
Number
[0288] In some embodiments, determining the initial first reliability score includes processing a third video frame (e.g., video frames 852, 854, and / or 856) among the plurality of video frames, where the third video frame corresponds to a third time (e.g., times T1, T2, and / or T3) before the first time.
[0289] In some embodiments, determining whether the user's visual attention is directed to the electronic device while the user is speaking includes using a machine learning model to determine whether the user's visual attention is directed to the electronic device while the user is speaking, including processing the representations of a first portion of the audio stream and a first portion of the video stream using the parameters of the machine learning model that represent the correlation between the user's mouth movements and the speech input.
[0290] In block 1106, in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking, a second portion of the audio stream is identified to include the user utterance addressed to the electronic device (e.g., by post-processing module 808).
[0291] In some embodiments, each of the plurality of reliability scores includes a fourth reliability score for a fourth video frame among the plurality of video frames (e.g., reliability score S4 for video frame 858). In some embodiments, identifying a second portion of the audio stream to include the user utterance addressed to the electronic device includes determining that a fourth time (e.g., time T4) corresponding to the fourth video frame is the start time of the second portion of the audio stream in accordance with the determination that the fourth reliability score exceeds a threshold.
[0292] In some embodiments, the plurality of video frames includes a fifth video frame and a sixth video frame consecutive to the fifth video frame. In some embodiments, identifying a second portion of the audio stream to include the user utterance addressed to the electronic device includes determining that a sixth time corresponding to the sixth video frame is the end time of the second portion of the audio stream in accordance with the determination that a fifth reliability score for the fifth video frame exceeds a second threshold and a sixth reliability score for the sixth video frame is below the second threshold.
[0293] In some embodiments, the second portion of the audio stream is identified to include the user utterance addressed to the electronic device without detecting an utterance trigger for starting a digital assistant session.
[0294] In block 1108, a task based on the second portion of the audio stream is initiated by a digital assistant operating on the electronic device, e.g., by digital assistant system 700.
[0295] In block 1110, an output indicating a started task is provided, for example, by device 900.
[0296] In some embodiments, the user is identified based on an audio stream (e.g., by identification module 810), and the output is based on the identified user. In some embodiments, the user is identified based on a video stream (e.g., by identification module 810), and the output is based on the identified user.
[0297] In block 1112, identifying a second portion of the audio stream to include user utterances addressed to the electronic device is discontinued according to a determination that the user's visual attention is not directed to the electronic device while the user is speaking.
[0298] In some embodiments, identifying a second portion of the audio stream to include user utterances addressed to the electronic device is discontinued according to a determination that the user's visual attention is directed to the electronic device while the user is not speaking.
[0299] In some embodiments, identifying a second portion of the audio stream to include user utterances addressed to the electronic device is discontinued according to a determination that the user's visual attention is not directed to the electronic device while the user is not speaking.
[0300] In some embodiments, identifying a second portion of the audio stream to include user utterances addressed to the electronic device is discontinued according to a determination that the user is not visible in a first portion of the video stream.
[0301] The operations described above with reference to FIG. 11 are optionally implemented by the components shown in FIGS. 1-4, 6A-6B, and 7A-7C and FIGS. 8A-8B. For example, the operations of process 1100 may be performed by digital assistant systems 700 and 800. Based on the components shown in FIGS. 1-4, 6A-6B, and 7A-7C, how other processes are implemented will be apparent to those skilled in the art.
[0302] FIG. 12 shows a process 1200 for determining whether a user's visual attention is directed towards an electronic device while the user is speaking, according to various examples. Process 1200 is performed, for example, using one or more electronic devices that implement a digital assistant. In some embodiments, process 1200 is performed using a client-server system (e.g., system 100), and the blocks of process 1200 are divided in any manner between a server (e.g., DA server 106) and a client device (e.g., device 900). In other embodiments, the blocks of process 1200 are divided between a server and multiple client devices (e.g., a mobile phone and a smartwatch). Thus, although some portions of process 1200 are described herein as being performed by certain devices of a client-server system, it will be understood that process 1200 is not so limited. In other embodiments, process 1200 is performed using only a client device (e.g., user device 104, device 900) or only multiple client devices. In process 1200, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in combination with process 1200.
[0303] In block 1202, an audio stream and a video stream are received simultaneously, for example, by device 900. In some embodiments, the audio stream includes speech from a first user and speech from a second user.
[0304] In block 1204, a first portion of the video stream is identified (e.g., by the preprocessing module 802) to include a first user, e.g., user 1000. In some embodiments, identifying the first portion of the video stream to include the first user includes tracking the head position of the first user. In some embodiments, identifying the first portion of the video stream includes cropping the video stream (e.g., by the preprocessing module 802) to obtain the first portion of the video stream.
[0305] In block 1206, a second portion of the video stream is identified (e.g., by the preprocessing module 802) to include a second user, e.g., user 1002. In some embodiments, identifying the second portion of the video stream to include the second user includes tracking the head position of the second user. In some embodiments, identifying the second portion of the video stream includes cropping the video stream (e.g., by the preprocessing module 802) to obtain the second portion of the video stream.
[0306] In block 1208, based on the first portion of the audio stream and the video stream, in accordance with a determination (e.g., by the audiovisual calling module 806) that the visual attention of the first user is directed towards the electronic device (e.g., device 900) while the first user is speaking, the first portion of the audio stream is identified (e.g., by the postprocessing module 808) to include the utterance of the first user addressed to the electronic device. In some embodiments, the first portion of the audio stream is identified to include the utterance of the first user addressed to the electronic device without detecting an utterance trigger for initiating a digital assistant session.
[0307] In some embodiments, a third portion of the audio stream received within a predetermined duration before the current time is stored in a memory buffer, such as memory buffer 804. In some embodiments, a first portion of the video stream is stored in a memory buffer, and the first portion of the video stream is received within a predetermined duration before the current time, and determining that the first user's visual attention is directed to the electronic device while the first user is speaking includes processing the stored content of the memory buffer.
[0308] In some embodiments, the first portion of the video stream includes a plurality of video frames (e.g., video frames 852, 854, 856, 858, 860, 862, and 864). In some embodiments, determining whether the first user's visual attention is directed to the electronic device while the first user is speaking includes determining an individual reliability score (e.g., reliability scores S1, S2, S3, S4, S5, S6, and S7) for each video frame of the plurality of video frames to obtain a plurality of respective reliability scores, and the individual reliability score indicates whether the first user's visual attention is directed to the electronic device while the first user is speaking for an individual video frame.
[0309] In some embodiments, the plurality of respective reliability scores includes a first reliability score (e.g., reliability score S4) for a first video frame (e.g., video frame 858) of the plurality of video frames, and the first video frame corresponds to a first time (e.g., time T4). In some embodiments, determining the first reliability score is based on the first video frame to obtain an initial first reliability score (e.g., initial reliability score
Number
[0310] In some embodiments, determining the initial first reliability score includes processing a third video frame (e.g., video frames 852, 854, and / or 856) among the plurality of video frames, wherein the third video frame corresponds to a third time (e.g., times T1, T2, and / or T3) before the first time.
[0311] In some embodiments, each of the plurality of reliability scores includes a fourth reliability score (e.g., S4) for a fourth video frame (e.g., video frame 858) among the plurality of video frames. In some embodiments, identifying a first portion of an audio stream as including speech of a first user addressed to an electronic device includes determining that a fourth time (e.g., time T4) corresponding to the fourth video frame is the start time of the first portion of the audio stream according to a determination that the fourth reliability score exceeds a threshold.
[0312] In some embodiments, the plurality of video frames includes a fifth video frame and a sixth video frame consecutive to the fifth video frame. In some embodiments, identifying a first portion of an audio stream as including speech of a first user addressed to an electronic device includes determining that a sixth time corresponding to the sixth video frame is the end time of the first portion of the audio stream according to a determination that a fifth reliability score for the fifth video frame exceeds a second threshold and a sixth reliability score for the sixth video frame is below the second threshold.
[0313] In block 1210, a first output is provided (e.g., by device 900) based on processing a first portion of the audio stream.
[0314] In some embodiments, a first user is identified (e.g., by identification module 810) based on at least one of a first portion of the video stream and the audio stream, and the first output is based on the identified first user. In some embodiments, providing the first output includes processing a first portion of the audio stream using a digital assistant (e.g., digital assistant system 700) operating on an electronic device.
[0315] In block 1212, according to a determination (e.g., by the audiovisual call module 806) that the visual attention of a second user is directed to the electronic device while the second user is speaking, based on the audio stream and a second portion of the video stream, the second portion of the audio stream is identified (e.g., by the post-processing module 808) to include the utterance of the second user addressed to the electronic device. In some embodiments, the second portion of the audio stream is identified to include the utterance of the second user addressed to the electronic device without detecting an utterance trigger.
[0316] In block 1214, a second output is provided (e.g., by device 900) based on processing the second portion of the audio stream.
[0317] In some embodiments, a second user is identified (e.g., by identification module 810) based on at least one of a second portion of the video stream and the audio stream, and the second output is based on the identified second user. In some embodiments, providing the second output includes processing the second portion of the audio stream using a digital assistant.
[0318] In some embodiments, based on the first portions of the audio stream and the video stream, in accordance with a determination (e.g., by the audiovisual call module 806) that the visual attention of the first user is directed to the electronic device while the first user is not speaking, identifying the first portion of the audio stream to include the utterance of the first user addressed to the electronic device is discontinued. In some embodiments, based on the second portions of the audio stream and the video stream, in accordance with a determination (e.g., by the audiovisual call module 806) that the visual attention of the second user is directed to the electronic device while the second user is not speaking, identifying the second portion of the audio stream to include the utterance of the second user addressed to the electronic device is discontinued.
[0319] In some embodiments, based on the first portions of the audio stream and the video stream, in accordance with a determination (e.g., by the audiovisual call module 806) that the visual attention of the first user is not directed to the electronic device while the first user is speaking, identifying the first portion of the audio stream to include the utterance of the first user addressed to the electronic device is discontinued. In some embodiments, based on the second portions of the audio stream and the video stream, in accordance with a determination (e.g., by the audiovisual call module 806) that the visual attention of the second user is not directed to the electronic device while the second user is speaking, identifying the second portion of the audio stream to include the utterance of the second user addressed to the electronic device is discontinued.
[0320] In some embodiments, based on the first portions of the audio stream and the video stream, in accordance with a determination (e.g., by the audiovisual call module 806) that the first user's visual attention is directed to the electronic device while the second user is speaking, identifying the first portion of the audio stream to include the first user's utterance addressed to the electronic device is aborted. In some embodiments, based on the second portions of the audio stream and the video stream, in accordance with a determination (e.g., by the audiovisual call module 806) that the second user's visual attention is directed to the electronic device while the first user is speaking, identifying the second portion of the audio stream to include the second user's utterance addressed to the electronic device is aborted.
[0321] The operations described above with reference to FIG. 12 may optionally be implemented by the components shown in FIGS. 1-4, FIGS. 6A-6B, and FIGS. 7A-7C and FIGS. 8A-8B. For example, the operations of process 1200 may be performed by digital assistant systems 700 and 800. Based on the components shown in FIGS. 1-4, FIGS. 6A-6B, and FIGS. 7A-7C, how other processes may be implemented will be apparent to those skilled in the art.
[0322] Figure 13 shows a process 1300 for determining whether a user's visual attention is directed towards an electronic device while the user is speaking, according to various examples. The process 1300 is executed using, for example, one or more electronic devices that implement a digital assistant. In some embodiments, the process 1300 is executed using a client-server system (e.g., system 100), and the blocks of the process 1300 are divided between a server (e.g., DA server 106) and a client device (e.g., device 900) in any manner. In other embodiments, the blocks of the process 1300 are divided between a server and a plurality of client devices (e.g., a mobile phone and a smartwatch). Thus, while some parts of the process 1300 are described herein as being executed by certain devices of a client-server system, it will be understood that the process 1300 is not so limited. In other embodiments, the process 1300 is executed using only a client device (e.g., user device 104, device 900) or only a plurality of client devices. In the process 1300, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in combination with the process 1300.
[0323] In block 1302, an audio stream and a video stream are received simultaneously, for example, by device 900.
[0324] In block 1304, a first type of reliability score (e.g., an audiovisual score) indicating whether a user's (e.g., user 904, 1000, or 1002) visual attention is directed towards an electronic device (e.g., device 900) while the user is speaking is determined based on the audio stream and the video stream (e.g., by audiovisual call module 806).
[0325] In some embodiments, determining the first type of reliability score includes determining the first type of reliability score based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within a predetermined duration before the current time.
[0326] In some embodiments, the first portion of the video stream includes a plurality of video frames (e.g., video frames 852, 854, 856, 858, 860, 862, and 864), and the first type of reliability score is for a first video frame (e.g., video frame 858) among the plurality of video frames. In some embodiments, the final reliability score is for the first video frame.
[0327] In block 1306, after determining the first type of reliability score, a final reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking is determined (e.g., by post - processing module 808) based on the first type of reliability score and a second type of reliability score indicating whether the user's visual attention is directed to the electronic device. In some embodiments, the second type of reliability score (e.g., a gaze score) indicates whether the user's line of sight is directed to the electronic device.
[0328] In some embodiments, determining the final reliability score includes determining the final reliability score based on a third type of reliability score (e.g., a posture score) indicating whether the user's posture is facing the electronic device.
[0329] In some embodiments, determining the final reliability score includes determining the final reliability score based on a fourth type of reliability score (e.g., a gesture score) indicating whether a predetermined type of user gesture is detected.
[0330] In some embodiments, determining the final reliability score includes determining the final reliability score based on a fifth type of reliability score (e.g., relative motion score) indicative of relative motion between the user and the electronic device.
[0331] In some embodiments, determining the final reliability score includes determining an action currently being performed by the electronic device, increasing the final reliability score in accordance with a determination that the action is of a first type, and decreasing the final reliability score in accordance with a determination that the action is of a second type different from the first type.
[0332] In some embodiments, determining the final reliability score includes determining the final reliability score based on data representative of user interactions with a digital assistant operating on the electronic device.
[0333] In block 1308, it is determined (e.g., by post - processing module 808) whether the final reliability score exceeds a threshold. In some embodiments, the threshold is decreased (e.g., by post - processing module 808) in accordance with a determination (e.g., by post - processing module 808) that the user is alone within the physical environment associated with the electronic device. In some embodiments, the threshold is decreased in accordance with a determination (e.g., by post - processing module 808) that the action currently being performed by the electronic device is of a first type. In some embodiments, the threshold is increased in accordance with a determination (e.g., by post - processing module 808) that the action currently being performed by the electronic device is of a second type different from the type. In some embodiments, the threshold is adjusted (e.g., by post - processing module 808) based on data representative of user interactions with a digital assistant operating on the electronic device.
[0334] In block 1310, in accordance with the determination that the final reliability score exceeds a threshold, a portion of the audio stream is identified (e.g., by post - processing module 808) to include user speech addressed to the electronic device. In some embodiments, identifying a portion of the audio stream to include user speech addressed to the electronic device includes determining that a first time (e.g., time T4) corresponding to a first video frame (e.g., video frame 858) is the start time of the portion of the audio stream. In some embodiments, a portion of the audio stream is identified to include user speech addressed to the electronic device without detecting a speech trigger to initiate a digital assistant session.
[0335] In some embodiments, identifying a portion of the audio stream to include user speech addressed to the electronic device includes determining that an initial start time of a portion of the audio stream is a second time corresponding to a first type of reliability score (e.g., a visual - auditory score), and adjusting the initial start time based on a third type of reliability score (e.g., a posture score).
[0336] In block 1312, based on processing a portion of the audio stream, an output is provided (e.g., by device 900). In some embodiments, providing the output includes processing a portion of the audio stream using a digital assistant operating on the electronic device, e.g., digital assistant system 700.
[0337] In block 1314, in accordance with the determination that the final reliability score does not exceed a threshold, identifying a portion of the audio stream that includes user speech addressed to the electronic device is aborted.
[0338] The operations described above with reference to FIG. 13 are optionally implemented by the components shown in FIGS. 1-4, 6A-6B, and 7A-7C and FIGS. 8A-8B. For example, the operations of process 1300 may be performed by digital assistant systems 700 and 800. Based on the components shown in FIGS. 1-4, 6A-6B, and 7A-7C, how other processes are implemented will be apparent to those skilled in the art.
[0339] According to some implementations, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided, which stores one or more programs executable by one or more processors of an electronic device, and the one or more programs include instructions for performing any of the methods or processes described herein.
[0340] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes means for performing any of the methods or processes described herein.
[0341] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes a processing unit configured to perform any of the methods or processes described herein.
[0342] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes one or more processors and a memory storing one or more programs for execution by the one or more processors, and the one or more programs include instructions for performing any of the methods or processes described herein.
[0343] The foregoing has been described with reference to specific embodiments for purposes of explanation. However, the above exemplary considerations are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. Embodiments have been selected and described in order to best explain the principles of the technology and their practical applications, thereby enabling others skilled in the art to best utilize the technology and various embodiments with various modifications as are suited to the particular use contemplated.
[0344] Although the present disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. Such changes and modifications are to be understood as being included within the scope of the present disclosure and examples as defined by the claims.
[0345] As described above, one aspect of the technology is to determine whether the visual attention of a user is directed to an electronic device while the user is speaking by collecting and using data available from various sources. The present disclosure contemplates that, in some examples, the collected data may include personal information data that uniquely identifies a particular person, or personal information data that can be used to contact a particular person or locate their whereabouts. Such personal information data can include demographic data, location-based data, phone numbers, email addresses, Twitter IDs, home addresses, data or records related to the user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.
[0346] The present disclosure recognizes that the use of such personal information data in the present technology can be a use that benefits the user. For example, personal information data can be used to provide relevant responses to user requests to a digital assistant. Thus, by using such personal information data, an electronic device can be used efficiently and accurately. Further, other uses related to personal information data that benefit the user are also contemplated by the present disclosure. For example, health data and fitness data can be used to provide insights into the user's overall wellness, or can also be used as positive feedback to individuals who are using technologies to pursue wellness goals.
[0347] This disclosure contemplates that entities involved in the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with firm privacy policies and / or privacy practices. Specifically, such entities should implement and consistently use privacy policies and practices that meet or exceed industry or government requirements for securely maintaining personal information data as confidential. Such policies should be readily accessible to users and updated as the data collection and / or use changes. Personal information from users should be collected for legitimate and proper use by the entity and should not be shared or sold except for those legitimate uses. Further, such collection / sharing should be done after informing and obtaining consent from the user. Moreover, such entities should consider taking all necessary measures to protect and secure access to such personal information data and ensure that others with access rights to personal information data faithfully adhere to their privacy policies and procedures. Additionally, such entities should be able to undergo third-party evaluations to demonstrate their compliance with widely accepted privacy policies and practices. Further, the policies and practices should be tailored to the specific types of personal information data being collected and / or accessed and should comply with applicable laws and regulations, including jurisdiction-specific considerations. For example, in the United States, the collection or access to certain health data may be subject to federal and / or state laws such as the Health Insurance Portability and Accountability Act (HIPAA). On the other hand, health data in other countries may be subject to different regulations and policies and should be addressed accordingly. Therefore, different privacy practices should be maintained for different types of personal data in each country.
[0348] Notwithstanding the foregoing, the present disclosure also contemplates embodiments in which a user can selectively block the use or access to personal information data. That is, the present disclosure is intended to provide hardware elements and / or software elements to prevent or block access to such personal information data. For example, when determining whether a user is looking and speaking, the technology can be configured to allow the user to select an "opt-in" or "opt-out" of participating in the collection of personal information data, either during registration of the service or at any time thereafter. In another example, the user can choose not to permit the use of personal information data to determine whether the user is looking and speaking. In yet another example, the user can choose to limit the amount of time that personal information data is collected to determine whether the user is looking and speaking.
[0349] In addition to providing the "opt-in" and "opt-out" options, the present disclosure is intended to provide notice regarding access to or use of personal information. For example, the user can be notified when downloading an application that will access the user's personal information data, and then again be alerted immediately before the personal information data is accessed by the application.
[0350] Furthermore, it is an intention of the present disclosure that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. The risk can be minimized by restricting the collection of data and deleting it when it is no longer needed. Additionally, anonymization of data can be used to protect the privacy of the user when applicable in certain health-related applications. Anonymization can be facilitated, as needed, by removing certain identifiers (e.g., date of birth, etc.), controlling the amount or specificity of the stored data (e.g., collecting location data at the city level rather than the address level), controlling how the data is stored (e.g., aggregating data across users), and / or other means.
[0351] Therefore, while the present disclosure broadly covers the use of personal information data for implementing one or more various disclosed embodiments, it is contemplated that the various embodiments can also be implemented without the need to access such personal information data. That is, the various embodiments of the present technology are not rendered inoperable by the absence of all or a portion of such personal information data. For example, a device can respond to voice input based on non-personal information data or a minimal amount of personal information, such as content requested by a device associated with the user, other non-personal information available to the device, or publicly available information.
Claims
Claim 1 A method comprising: In an electronic device having one or more processors and a memory, Receiving an audio stream and a video stream simultaneously; Based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within the predetermined duration before the current time, determining whether the user's visual attention is directed towards the electronic device while the user is speaking; While the user is speaking, in accordance with the determination that the user's visual attention is directed towards the electronic device, Identifying a second portion of the audio stream to include user utterances addressed to the electronic device; Initiating a task based on the second portion of the audio stream by a digital assistant operating on the electronic device; Providing an output indicating the initiated task; A method comprising the above steps. Claim 2 Determining whether the user's visual attention is directed towards the electronic device while the user is speaking includes: Determining whether the user's visual attention is directed towards the display of the electronic device. The method according to claim 1. Claim 3 Determining whether the user's visual attention is directed towards the electronic device while the user is speaking includes: Determining whether the user's visual attention is directed towards an affordance displayed by the electronic device. The method according to claim 1. Claim 4 The second portion of the audio stream is identified to include user utterances addressed to the electronic device without detecting an utterance trigger for starting a digital assistant session. The method according to any one of claims 1 to 3. Claim 5 The method further includes storing the first portion of the audio stream and the first portion of the video stream in a memory buffer. Determining whether the user's visual attention is directed towards the electronic device while the user is speaking includes processing the stored content in the memory buffer. The method according to any one of claims 1 to 4. Claim 6 The first portion of the video stream includes a plurality of video frames, determining whether the visual attention of the user is directed towards the electronic device while the user is speaking, includes determining an individual reliability score for each video frame of the plurality of video frames to obtain a plurality of respective reliability scores, the individual reliability score indicating whether the visual attention of the user is directed towards the electronic device while the user is speaking for the individual video frame, the method according to any one of claims 1 to 5.
7. The plurality of respective reliability scores includes a first reliability score for a first video frame of the plurality of video frames, the first video frame corresponds to a first time, determining the first reliability score, includes determining an initial first reliability score based on the first video frame, and adjusting the initial first reliability score to obtain the first reliability score based on processing a second video frame of the plurality of video frames, the second video frame corresponding to a second time after the first time, the method according to claim 6.
8. Determining the initial first reliability score includes processing a third video frame of the plurality of video frames, the third video frame corresponding to a third time before the first time, the method according to claim 7.
9. The plurality of respective reliability scores includes a fourth reliability score for a fourth video frame of the plurality of video frames, identifying the second portion of the audio stream to include user speech addressed to the electronic device, in accordance with a determination that the fourth reliability score exceeds a threshold, includes determining that the fourth time corresponding to the fourth video frame is the start time of the second portion of the audio stream, the method according to any one of claims 6 to 8.
10. The plurality of video frames includes a fifth video frame and a sixth video frame consecutive to the fifth video frame, identifying the second portion of the audio stream to include user speech addressed to the electronic device, according to a determination that a fifth reliability score for the fifth video frame exceeds a second threshold and that a sixth reliability score for the sixth video frame is below the second threshold, the method according to any one of claims 6 to 9, comprising determining that a sixth time corresponding to the sixth video frame is an end time of the second portion of the audio stream.
11. determining whether the user's visual attention is directed to the electronic device while the user is speaking, using a machine learning model to determine whether the user's visual attention is directed to the electronic device while the user is speaking, processing representations of the first portion of the audio stream and the first portion of the video stream using parameters of the machine learning model that represent a correlation between the user's mouth movements and speech input, the method according to any one of claims 1 to 10.
12. determining that the user's visual attention is directed to the electronic device while the user is speaking, the method according to any one of claims 1 to 11, comprising determining that the user's line of sight is directed to the electronic device while the user is speaking.
13. determining that the user's visual attention is directed to the electronic device while the user is speaking, the method according to any one of claims 1 to 12, comprising determining that the user is facing the electronic device while the user is speaking.
14. the method according to any one of claims 1 to 13, further comprising identifying the user based on the audio stream, wherein the output is based on the identified user.
15. the method according to any one of claims 1 to 14, further comprising identifying the user based on the video stream, wherein the output is based on the identified user.
16. according to a determination that the user's visual attention is not directed to the electronic device while the user is speaking, The method according to any one of claims 1 to 15, further comprising ceasing to identify the second portion of the audio stream to include user speech addressed to the electronic device.
17. According to a determination that the user's visual attention is directed to the electronic device while the user is not speaking, The method according to any one of claims 1 to 16, further comprising ceasing to identify the second portion of the audio stream to include user speech addressed to the electronic device.
18. According to a determination that the user's visual attention is not directed to the electronic device while the user is not speaking, The method according to any one of claims 1 to 17, further comprising ceasing to identify the second portion of the audio stream to include user speech addressed to the electronic device.
19. In the first portion of the video stream, according to a determination that the user is not visible, The method according to any one of claims 1 to 18, further comprising ceasing to identify the second portion of the audio stream to include user speech addressed to the electronic device.
20. An electronic device, One or more processors, A memory, One or more programs, an electronic device comprising, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions being, Receiving an audio stream and a video stream simultaneously, Based on the first portion of the audio stream received within a predetermined duration before the current time and the first portion of the video stream received within the predetermined duration before the current time, determining whether the user's visual attention is directed to the electronic device while the user is speaking, While the user is speaking, according to a determination that the user's visual attention is directed to the electronic device, Identifying a second portion of the audio stream to include user speech addressed to the electronic device, Starting a task based on the second portion of the audio stream by a digital assistant operating on the electronic device, An electronic device that provides an output indicating the started task.
21. A non - transitory computer - readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive an audio stream and a video stream simultaneously, determine whether the user's visual attention is directed to the electronic device while the user is speaking, based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within the predetermined duration before the current time, in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking, identify a second portion of the audio stream to include a user utterance addressed to the electronic device, cause a digital assistant operating on the electronic device to start a task based on the second portion of the audio stream, provide an output indicating the started task. A non - transitory computer - readable storage medium.
22. An electronic device, comprising means for receiving an audio stream and a video stream simultaneously, means for determining whether the user's visual attention is directed to the electronic device while the user is speaking, based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within the predetermined duration before the current time, in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking, means for identifying a second portion of the audio stream to include a user utterance addressed to the electronic device, means for causing a digital assistant operating on the electronic device to start a task based on the second portion of the audio stream, means for providing an output indicating the started task. An electronic device comprising the above.
23. An electronic device, comprising one or more processors, a memory, An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 19.
24. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 19.
25. An electronic device, comprising means for executing the method according to any one of claims 1 to 19. Electronic device.
26. A method, in an electronic device having one or more processors and a memory, receiving an audio stream and a video stream simultaneously; identifying a first portion of the video stream that includes a first user; identifying a second portion of the video stream that includes a second user; based on the audio stream and the first portion of the video stream, in accordance with a determination that the visual attention of the first user is directed to the electronic device while the first user is speaking, identifying a first portion of the audio stream that includes the utterance of the first user addressed to the electronic device; providing a first output based on processing the first portion of the audio stream; based on the audio stream and the second portion of the video stream, in accordance with a determination that the visual attention of the second user is directed to the electronic device while the second user is speaking, identifying a second portion of the audio stream that includes the utterance of the second user addressed to the electronic device; providing a second output based on processing the second portion of the audio stream; comprising.
27. Identifying the first portion of the video stream that includes the first user includes tracking the head position of the first user. Identifying the second portion of the video stream to include the second user includes tracking the head position of the second user, the method of claim 26. **Claim 28** Identifying the first portion of the video stream includes cropping the video stream to obtain the first portion of the video stream. Identifying the second portion of the video stream includes cropping the video stream to obtain the second portion of the video stream, the method of claim 26 or 27. **Claim 29** Identifying the first user based on at least one of the first portion of the video stream and the audio stream, wherein the first output is based on the first user to be identified. Identifying the second user based on at least one of the second portion of the video stream and the audio stream, wherein the second output is based on the second user to be identified, the method of any one of claims 26 to 28. **Claim 30** The audio stream includes speech of the first user and speech of the second user, the method of any one of claims 26 to 29. **Claim 31** Based on the audio stream and the first portion of the video stream, according to the determination that the visual attention of the first user is directed to the electronic device while the first user is not speaking. Ceasing to identify the first portion of the audio stream that includes speech of the first user addressed to the electronic device. Based on the audio stream and the second portion of the video stream, according to the determination that the visual attention of the second user is directed to the electronic device while the second user is not speaking. Ceasing to identify the second portion of the audio stream to include speech of the second user addressed to the electronic device, the method of any one of claims 26 to 30. **Claim 32** Based on the first portions of the audio stream and the video stream, in accordance with the determination that the visual attention of the first user is not directed towards the electronic device while the first user is speaking, ceasing to identify the first portion of the audio stream that includes the utterance of the first user addressed to the electronic device, Based on the second portions of the audio stream and the video stream, in accordance with the determination that the visual attention of the second user is not directed towards the electronic device while the second user is speaking, ceasing to identify the second portion of the audio stream so as to include the utterance of the second user addressed to the electronic device, the method according to any one of claims 26 to 31, further comprising.
33. Based on the first portions of the audio stream and the video stream, in accordance with the determination that the visual attention of the first user is directed towards the electronic device while the second user is speaking, ceasing to identify the first portion of the audio stream that includes the utterance of the first user addressed to the electronic device, Based on the second portions of the audio stream and the video stream, in accordance with the determination that the visual attention of the second user is directed towards the electronic device while the first user is speaking, ceasing to identify the second portion of the audio stream so as to include the utterance of the second user addressed to the electronic device, the method according to any one of claims 26 to 32, further comprising.
34. storing a third portion of the audio stream received within a predetermined duration prior to the current time in a memory buffer, storing the first portion of the video stream in the memory buffer, the first portion of the video stream being received within the predetermined duration prior to the current time, and determining that the visual attention of the first user is directed towards the electronic device while the first user is speaking includes processing the stored content of the memory buffer, the method according to any one of claims 26 to 33.
35. The first portion of the video stream includes a plurality of video frames, Determining that the visual attention of the first user is directed towards the electronic device while the first user is speaking, Including determining an individual reliability score for each video frame among the plurality of video frames to obtain a plurality of respective reliability scores, the individual reliability score indicating whether the visual attention of the first user is directed towards the electronic device while the first user is speaking for the individual video frame, the method according to any one of claims 26 to 34.
36. The plurality of respective reliability scores includes a first reliability score for a first video frame among the plurality of video frames, The first video frame corresponds to a first time, Determining the first reliability score includes: Determining an initial first reliability score based on the first video frame; and Adjusting the initial first reliability score to obtain the first reliability score based on processing a second video frame among the plurality of video frames, the second video frame corresponding to a second time after the first time, the method according to claim 35.
37. Determining the initial first reliability score includes processing a third video frame among the plurality of video frames, the third video frame corresponding to a third time before the first time, the method according to claim 36.
38. The plurality of respective reliability scores includes a fourth reliability score for a fourth video frame among the plurality of video frames, Identifying the first portion of the audio stream including the utterance of the first user addressed to the electronic device, According to the determination that the fourth reliability score exceeds a threshold, Including determining that the fourth time corresponding to the fourth video frame is the start time of the first portion of the audio stream, the method according to any one of claims 35 to 37.
39. The plurality of video frames includes a fifth video frame and a sixth video frame consecutive to the fifth video frame, Identifying the first portion of the audio stream that includes speech of the first user addressed to the electronic device, According to a determination that a fifth reliability score for the fifth video frame exceeds a second threshold and a sixth reliability score for the sixth video frame is below the second threshold, Determining that a sixth time corresponding to the sixth video frame is an end time of the first portion of the audio stream, the method according to any one of claims 35 to 38.
40. The first portion of the audio stream is identified to include speech of the first user addressed to the electronic device without detecting a speech trigger for starting a digital assistant session, The second portion of the audio stream is identified to include speech of the second user addressed to the electronic device without detecting the speech trigger, the method according to any one of claims 26 to 39.
41. Providing the first output includes processing the first portion of the audio stream using a digital assistant operating on the electronic device, Providing the second output includes processing the second portion of the audio stream using the digital assistant, the method according to any one of claims 26 to 40.
42. An electronic device, One or more processors, A memory, One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs include instructions, and the instructions Receive an audio stream and a video stream simultaneously, Identify a first portion of the video stream that includes a first user, Identify a second portion of the video stream that includes a second user, Based on the first portion of the audio stream and the first portion of the video stream, according to a determination that the visual attention of the first user is directed to the electronic device while the first user is speaking, Identify a first portion of the audio stream that includes speech of the first user addressed to the electronic device, Based on processing the first portion of the audio stream, provide a first output. Based on the second portion of the audio stream and the video stream, according to the determination that the visual attention of the second user is directed towards the electronic device while the second user is speaking. Identify a second portion of the audio stream to include the utterance of the second user addressed to the electronic device. An electronic device that provides a second output based on processing the second portion of the audio stream. **Claim 43** A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to Receive an audio stream and a video stream simultaneously. Identify a first portion of the video stream to include a first user. Identify a second portion of the video stream to include a second user. Based on the first portion of the audio stream and the video stream, according to the determination that the visual attention of the first user is directed towards the electronic device while the first user is speaking. Identify a first portion of the audio stream that includes the utterance of the first user addressed to the electronic device. Provide a first output based on processing the first portion of the audio stream. Based on the second portion of the audio stream and the video stream, according to the determination that the visual attention of the second user is directed towards the electronic device while the second user is speaking. Identify a second portion of the audio stream to include the utterance of the second user addressed to the electronic device. A non-transitory computer-readable storage medium that causes a second output to be provided based on processing the second portion of the audio stream. **Claim 44** An electronic device, comprising Receiving an audio stream and a video stream simultaneously. Identifying a first portion of the video stream to include a first user. Identifying a second portion of the video stream to include a second user. Based on the first part of the audio stream and the video stream, in accordance with the determination that the visual attention of the first user is directed towards the electronic device while the first user is speaking, identifying a first part of the audio stream that includes the utterance of the first user addressed to the electronic device, means for providing a first output based on processing the first part of the audio stream, Based on the second part of the audio stream and the video stream, in accordance with the determination that the visual attention of the second user is directed towards the electronic device while the second user is speaking, identifying a second part of the audio stream that includes the utterance of the second user addressed to the electronic device, means for providing a second output based on processing the second part of the audio stream, An electronic device comprising the above.
45. An electronic device, one or more processors, a memory, one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 26 to 41. An electronic device.
46. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to execute the method according to any one of claims 26 to 41. A non-transitory computer-readable storage medium.
47. An electronic device, comprising means for executing the method according to any one of claims 26 to 41, An electronic device.
48. A method, in an electronic device having one or more processors and a memory, simultaneously receiving an audio stream and a video stream, determining a first type of reliability score indicating whether the visual attention of the user is directed towards the electronic device while the user is speaking, based on the audio stream and the video stream, after determining the first type of reliability score, Determining a final reliability score indicating whether the user's visual attention is directed towards the electronic device while the user is speaking, based on the reliability score of the first type and a reliability score of a second type indicating whether the user's visual attention is directed towards the electronic device; Identifying a portion of the audio stream that includes user speech addressed to the electronic device according to a determination that the final reliability score exceeds a threshold; Providing an output based on processing the portion of the audio stream; A method comprising.
49. Determining the reliability score of the first type includes Determining the reliability score of the first type based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within the predetermined duration before the current time, the method according to claim 48.
50. The first portion of the video stream includes a plurality of video frames, The reliability score of the first type is for a first video frame among the plurality of video frames, the method according to claim 49.
51. The final reliability score is for the first video frame, Identifying the portion of the audio stream that includes user speech addressed to the electronic device includes Determining that a first time corresponding to the first video frame is a start time of the portion of the audio stream, the method according to claim 50.
52. The reliability score of the second type indicates whether the user's line of sight is directed towards the electronic device, the method according to any one of claims 48 to 51.
53. Determining the final reliability score includes Determining the final reliability score based on a reliability score of a third type indicating whether the user's posture is facing the electronic device, the method according to any one of claims 48 to 52.
54. Identifying the portion of the audio stream that includes user speech addressed to the electronic device includes Determining that the initial start time of the part of the audio stream is a second time corresponding to the reliability score of the first type; Adjusting the initial start time based on the reliability score of the third type, the method according to claim 53. **Claim 55** Determining the final reliability score includes: Determining the final reliability score based on a fourth type of reliability score indicating whether a gesture of a user of a predetermined type is detected, the method according to any one of claims 48 to 54. **Claim 56** Determining the final reliability score includes: Determining the final reliability score based on a fifth type of reliability score indicating the relative movement between the user and the electronic device, the method according to any one of claims 48 to 55. **Claim 57** According to the determination that the user is alone in the physical environment associated with the electronic device, Further including reducing the threshold value, the method according to any one of claims 48 to 56. **Claim 58** Determining the final reliability score includes: Determining the action currently being performed by the electronic device; According to the determination that the action is of a first type, Increasing the final reliability score; According to the determination that the action is of a second type different from the first type, Reducing the final reliability score, the method according to any one of claims 48 to 57. **Claim 59** According to the determination that the action currently being performed by the electronic device is of a first type, Reducing the threshold value; According to the determination that the action currently being performed by the electronic device is of a second type different from the type, Further including increasing the threshold value, the method according to any one of claims 48 to 58. **Claim 60** Determining the final reliability score includes: Determining the final reliability score based on data representing a user interaction with a digital assistant operating on the electronic device, the method according to any one of claims 48 to 59. **Claim 61** The method according to any one of claims 48 to 60, further comprising adjusting the threshold based on data representing a user interaction with a digital assistant operating on the electronic device.
62. The method according to any one of claims 48 to 61, wherein providing the output includes processing the portion of the audio stream using a digital assistant operating on the electronic device.
63. The method according to any one of claims 48 to 62, wherein the portion of the audio stream is identified to include a user utterance addressed to the electronic device without detecting an utterance trigger for starting a digital assistant session.
64. In accordance with the determination that the final reliability score does not exceed the threshold, The method according to any one of claims 48 to 63, further comprising ceasing to identify the portion of the audio stream to include a user utterance addressed to the electronic device.
65. An electronic device, comprising: One or more processors; A memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions, and the instructions are: Receiving an audio stream and a video stream simultaneously; Based on the audio stream and the video stream, determining a first type of reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking; After determining the first type of reliability score, Based on the first type of reliability score and a second type of reliability score indicating whether the user's visual attention is directed to the electronic device, determining a final reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking; In accordance with the determination that the final reliability score exceeds a threshold, Identifying a portion of the audio stream to include a user utterance addressed to the electronic device; An electronic device that provides an output based on processing the portion of the audio stream.
66. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive an audio stream and a video stream simultaneously, based on the audio stream and the video stream, determine a first type of reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking, after determining the first type of reliability score, based on the first type of reliability score and a second type of reliability score indicating whether the user's visual attention is directed to the electronic device, determine a final reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking, in accordance with a determination that the final reliability score exceeds a threshold, identify a portion of the audio stream that includes user utterances addressed to the electronic device, and provide an output based on processing the portion of the audio stream. A non-transitory computer-readable storage medium. **Claim 67** An electronic device, receiving an audio stream and a video stream simultaneously, based on the audio stream and the video stream, determining a first type of reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking, after determining the first type of reliability score, means for determining a final reliability score indicating whether the user's visual attention is directed to the electronic device while the user is speaking, based on the first type of reliability score and a second type of reliability score indicating whether the user's visual attention is directed to the electronic device, in accordance with a determination that the final reliability score exceeds a threshold, identifying a portion of the audio stream that includes user utterances addressed to the electronic device, and means for providing an output based on processing the portion of the audio stream. An electronic device comprising the above. **Claim 68** An electronic device, one or more processors, a memory, An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 48 to 64.
69. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to execute the method according to any one of claims 48 to 64.
70. An electronic device, comprising means for executing the method according to any one of claims 48 to 64. Electronic device.
71. A computer program product comprising one or more programs configured to be executed by one or more processors of an electronic device, wherein the one or more programs include instructions that simultaneously receive an audio stream and a video stream, determine whether the user's visual attention is directed to the electronic device while the user is speaking, based on a first portion of the audio stream received within a predetermined duration before the current time and a first portion of the video stream received within the predetermined duration before the current time, in accordance with the determination that the user's visual attention is directed to the electronic device while the user is speaking, identify a second portion of the audio stream that includes a user utterance addressed to the electronic device, initiate a task based on the second portion of the audio stream by a digital assistant operating on the electronic device, and provide an output indicating the initiated task.
72. A computer program product including one or more programs configured to be executed by one or more processors of an electronic device, wherein the one or more programs include instructions for executing the method according to any one of claims 1 to 19.
73. A computer program product comprising one or more programs configured to be executed by one or more processors of an electronic device, wherein the one or more programs include instructions, and the instructions are simultaneously receive an audio stream and a video stream, identify a first portion of the video stream that includes a first user, identify a second portion of the video stream that includes a second user, based on the audio stream and the first portion of the video stream, in accordance with a determination that the visual attention of the first user is directed to the electronic device while the first user is speaking, identify a first portion of the audio stream that includes the utterance of the first user addressed to the electronic device, provide a first output based on processing the first portion of the audio stream, based on the audio stream and the second portion of the video stream, in accordance with a determination that the visual attention of the second user is directed to the electronic device while the second user is speaking, identify a second portion of the audio stream that includes the utterance of the second user addressed to the electronic device, provide a second output based on processing the second portion of the audio stream. A computer program product.
74. A computer program product including one or more programs configured to be executed by one or more processors of an electronic device, wherein the one or more programs include instructions for executing the method according to any one of claims 26 to 41. A computer program product.
75. A computer program product comprising one or more programs configured to be executed by one or more processors of an electronic device, wherein the one or more programs include instructions, and the instructions are simultaneously receive an audio stream and a video stream, based on the audio stream and the video stream, determine a first type of reliability score indicating whether the visual attention of the user is directed to the electronic device while the user is speaking, after determining the first type of reliability score, Based on the first type of reliability score and a second type of reliability score indicating whether the user's visual attention is directed towards the electronic device, determine a final reliability score indicating whether the user's visual attention is directed towards the electronic device while the user is speaking. In accordance with a determination that the final reliability score exceeds a threshold, Identify a portion of the audio stream that includes user speech addressed to the electronic device. A computer program product that provides an output based on processing the portion of the audio stream. **Claim 76** A computer program product including one or more programs configured to be executed by one or more processors of an electronic device, the one or more programs including instructions to perform the method according to any one of claims 48 to 64.
Citation Information
Patent Citations
Managing agent engagement in a man-machine dialog
JP2018180523A
Information processing apparatus and method, and program
JP2021144259A
Adapting an automated assistant based on detected mouth movements and / or gaze
JP2021521497A
Sensor Fusion Service to Enhance Human Computer Interactions
US20190138268A1