Voice Triggers for Digital Assistants

The low-power voice trigger system addresses the challenge of initiating voice-based digital assistants without haptic input and reduces power consumption, enabling an efficient and hands-free user experience.

JP7689644B1Active Publication Date: 2025-06-06APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025033999
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-02-07
Filing Date
2025-03-04
Publication Date
2025-06-06
Estimated Expiration
2034-02-07

AI Technical Summary

Technical Problem

Existing voice-based digital assistants require haptic input to initiate, which inhibits the hands-free experience and consumes significant power when continuously listening for voice input.

Method used

A low-power voice trigger system that uses a combination of sound detectors, including a noise detector, sound type detector, and trigger sound detector, to activate a voice-based digital assistant without the need for haptic input, while minimizing power consumption.

Benefits of technology

Enables an 'always listening' voice trigger function with reduced power consumption, allowing users to initiate voice-based services hands-free while conserving battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689644000001
    Figure 0007689644000001
  • Figure 0007689644000002
    Figure 0007689644000002
  • Figure 0007689644000003
    Figure 0007689644000003
Patent Text Reader

Abstract

A method for operating a voice trigger is provided. In some implementations, the method is performed in an electronic device that includes one or more processors and a memory that stores instructions executed by the one or more processors. The method includes receiving an audio input. The audio input may correspond to a spoken word or phrase, or a portion thereof. The method includes determining whether at least a portion of the audio input corresponds to a predetermined type of sound, such as a human voice. The method includes, upon determining that at least a portion of the audio input corresponds to the predetermined type, determining whether the audio input includes predetermined content, such as a predetermined trigger word or phrase. The method also includes, upon determining that the audio input includes the predetermined content, initiating a speech-based service, such as a voice-based digital assistant.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 61 / 762,260, entitled "VOICE TRIGGER FOR A DIGITAL ASSISTANT," filed February 7, 2013, and is hereby incorporated by reference in its entirety for all purposes.

[0002] <Technical field> The disclosed embodiments relate generally to digital assistants, and more particularly, to methods and systems for voice triggers for digital assistants. [Background technology]

[0003] Recently, voice-based digital assistants, such as Apple's SIRI®, have been introduced to the market to handle various tasks, such as searching and navigating the web. One advantage of such voice-based digital assistants is that a user can interact with the device in a hands-free manner without manipulating or looking at the device. Hands-free operation can be particularly useful when a person cannot or should not physically operate the device, such as while driving. However, to initiate a voice-based assistant, a user typically needs to press a button or select an icon on a touch screen. This haptic input inhibits the hands-free experience. Accordingly, it would be advantageous to provide a method and system for enabling a voice-based digital assistant (or other speech-based service) using voice input or signals rather than haptic input.

[0004] Enabling a voice-based assistant using voice input requires monitoring an audio channel to detect voice input. This monitoring consumes power, which is a limited resource on the battery-dependent handheld or portable devices on which such voice-based digital assistants often run. It would therefore be beneficial to provide an energy-efficient voice trigger that can be used to initiate voice-based and / or speech-based services on the device. Summary of the Invention

[0005] Accordingly, there is a need for a low-power voice trigger that can provide an "always listening" voice trigger function without excessively consuming limited power resources. The embodiments described below provide a system and method for initiating a voice-based assistant using a voice trigger on an electronic device. Interaction with a voice-based digital assistant (or other speech-based service, such as a speech-to-text transcription service) often begins when a user presses an affordance (e.g., a button or icon) on the device to activate the digital assistant, and the device then provides some indication to the user that the digital assistant is active and listening, such as a light, a sound (e.g., a beep), or a spoken output (e.g., "What can I do for you?"). As described herein, voice triggers can also be implemented to be activated in response to a specific and predefined word, phrase, or sound, without requiring physical interaction by the user. For example, a user may be able to activate the SIRI digital assistant on an IPHONE (both of which are provided by Apple Inc., the assignee of the present application) by speaking the phrase "call SIRI." In response, the device may emit a beep, tone, or speech output (e.g., "What can I do for you?") to indicate to the user that listening mode is active. In response, the user may begin interacting with the digital assistant without having to physically touch the device that provides the digital assistant functionality.

[0006] One technique for initiating a speech-based service with a voice trigger is to have the speech-based service continuously listen for a predefined trigger word, phrase, or sound (any of which may be referred to herein as a "trigger sound"). However, continuously operating a speech-based service (e.g., a voice-based digital assistant) requires significant voice processing and battery power. To reduce power consumption in providing a voice trigger function, various techniques may be employed. In some implementations, the main processor (i.e., the "application processor") of the electronic device is maintained in a low-power or no-power state while one or more sound detectors of low power consumption (e.g., to be independent of the application processor) are maintained active. (When in a low-power or no-power state, the application processor or any other processor, program, or module may be described as being in a disabled or standby mode.) For example, even when the application processor is disabled, a low-power sound detector is used to monitor the audio channel for a trigger sound. This sound detector is sometimes referred to herein as a trigger sound detector. In some implementations, it is configured to detect a particular sound, phoneme, and / or word. Trigger sound detectors (including hardware and / or software components) are designed to recognize distinctive words, sounds, or phrases, but are generally not capable of or optimized to provide full speech-to-text functionality, as such tasks require significant computational and power resources. Thus, in some implementations, the trigger sound detector recognizes whether the voice input contains a predefined pattern (e.g., a sonic pattern that matches the words "to SIRI"), but is not capable of (or configured to) convert the voice input to text or recognize many other words. When a trigger sound is detected, the digital assistant is subsequently taken out of standby mode so that the user can provide a voice command.

[0007] In some implementations, the trigger sound detector is configured to detect a variety of different trigger sounds, such as a set of words, phrases, sounds, and / or combinations thereof. The user can then use any of these sounds to initiate a speech-based service. In one example, the voice trigger is preconfigured to respond to the phrases "Call SIRI", "SIRI launch", "Call digital assistant", or "Hello, HAL, can you hear me, HAL?". In some implementations, the user must select one of the preconfigured trigger sounds as the single trigger sound. In some implementations, the user selects a subset of the preconfigured trigger sounds so that the user can initiate a speech-based service with different trigger sounds. In some implementations, all of the preconfigured trigger sounds remain valid trigger sounds.

[0008] In some implementations, a separate sound detector is used, and the trigger sound detector may also be maintained in a low or no power mode much of the time. For example, a different type of sound detector (e.g., one that uses less power than the trigger sound detector) is used to monitor the audio channel and determine if the sound input corresponds to a particular type of sound. Sounds are classified into different "types" based on certain identifiable sound characteristics. For example, sounds that belong to the type "human voice" have a particular spectral content, periodicity, fundamental frequency, etc. Other types of sounds (e.g., whistling, clapping, etc.) have different characteristics. The different types of sounds are identified using voice processing and / or signal processing techniques, as described herein. This sound detector is sometimes referred to herein as a "sound type detector." For example, if the predefined trigger phrase is "To SIRI," the sound type detector determines if the input approximately corresponds to a human speaking voice. If the trigger sound is a non-voiced sound, such as a whistle, the sound type detector determines if the sound input approximately corresponds to a whistle. When an appropriate type of sound is detected, the sound type detector activates the trigger sound detector to further process and / or analyze the sound. Because the sound type detector requires less power than the trigger sound detector (e.g., because it uses lower power requirements circuitry and / or more efficient audio processing algorithms than the trigger sound detector), the voice trigger function consumes less power than the trigger sound detector alone.

[0009] In some implementations, an additional sound detector is used, and both the sound type detector and the trigger sound detector may be kept in a low or no power mode much of the time. For example, a sound detector, using less power than the sound type detector, may be used to monitor an audio channel to determine if the sound input meets a predetermined condition, such as an amplitude threshold (e.g., volume). This sound detector may also be referred to herein as a noise detector. When the noise detector detects a sound that meets a predetermined threshold, the noise detector activates the sound type detector to further process and / or analyze the sound. Because the noise detector requires less power than the sound type detector or the trigger sound detector (e.g., because it uses less power-demanding circuitry and / or more efficient audio processing algorithms), the voice trigger function consumes less power than the combination of the sound type detector and the trigger sound detector without the noise detector.

[0010] In some implementations, any one or more of the sound detectors described above are operated according to a duty cycle that cycles between "on" and "off" states. This further helps reduce the power consumption of the voice trigger. For example, in some implementations, the noise detector is "on" (i.e., actively monitoring the audio channel) for 10 milliseconds, followed by "off" for 90 milliseconds. In this way, the noise detector is "off" 90% of the time while still effectively providing a continuous noise detection function. In some implementations, the on and off durations for the sound detectors are selected such that all of the detectors are enabled while the trigger sound is still being input. For example, for a trigger phrase "to SIRI", the sound detectors may be configured such that wherever the trigger phrase begins in the duty cycle(s), the trigger sound detector is enabled in time to analyze a sufficient amount of input. For example, the trigger sound detector is enabled in time to receive, process, and analyze enough of the sound "to IRI" to determine that the sound matches the trigger phrase. In some implementations, the sound input is stored in memory as it is received and passed to an upstream detector so that a majority of the sound input can be analyzed. Accordingly, even if the trigger sound detector is not activated until after the trigger phrase is spoken, the entirety of the recorded trigger phrase can still be analyzed.

[0011] Some embodiments provide a method of operating a voice trigger. The method is implemented in an electronic device that includes one or more processors and a memory that stores instructions executed by the one or more processors. The method includes receiving a sound input. The method further includes determining whether at least a portion of the sound input corresponds to a predetermined type of sound. The method further includes determining whether the sound input includes predetermined content upon determining that at least a portion of the sound input corresponds to the predetermined type. The method further includes initiating a speech-based service upon determining that the sound input includes the predetermined content. In some embodiments, the speech-based service is a voice-based digital assistant. In some embodiments, the speech-based service is a dictation service.

[0012] In some implementations, determining whether the sound input corresponds to a predetermined type of sound is performed by a first sound detector and determining whether the sound input includes predetermined content is performed by a second sound detector. In some implementations, the first sound detector consumes less power in operation than the second sound detector. In some implementations, the first sound detector performs a frequency domain analysis of the sound input. In some implementations, determining whether the sound input corresponds to a predetermined type of sound is performed upon determining that the sound input satisfies a predetermined condition (e.g., as determined by a third sound detector, described below).

[0013] In some implementations, the first sound detector periodically monitors the audio channel according to a duty cycle, which in some implementations includes an on-time of about 20 milliseconds and an off-time of about 100 milliseconds.

[0014] In some implementations, the predetermined type is a human voice and the predetermined content is one or more words. In some implementations, determining whether at least a portion of the sound input corresponds to the predetermined type of sound includes determining whether at least a portion of the sound input includes a frequency characteristic of a human voice.

[0015] In some implementations, the second sound detector is activated in response to the first sound detector determining that the sound input corresponds to a predetermined type. In some implementations, the second sound detector is operated for at least a predetermined time period after the first sound detector determines that the sound input corresponds to a predetermined type. In some implementations, the predetermined time period corresponds to a duration of a predetermined content.

[0016] In some implementations, the predetermined content is one or more predetermined phonemes. In some implementations, the one or more predetermined phonemes comprise at least one word.

[0017] In some implementations, the method includes determining whether the sound input satisfies a predetermined condition prior to determining whether the sound input corresponds to a predetermined type of sound. In some implementations, the predetermined condition is an amplitude threshold. In some implementations, determining whether the sound input satisfies the predetermined condition is performed by a third sound detector, the third sound detector consuming less power in operation than the first sound detector. In some implementations, the third sound detector periodically monitors the audio channel according to a duty cycle. In some implementations, the duty cycle includes an on-time of about 20 milliseconds and an off-time of about 500 milliseconds. In some implementations, the third sound detector performs a time domain analysis of the sound input.

[0018] In some implementations, the method includes storing at least a portion of the sound input in a memory, and providing the portion of the sound input to the speech-based service when the speech-based service is initiated. In some implementations, the portion of the sound input is stored in the memory using direct memory access.

[0019] In some embodiments, the method includes determining whether the sound input corresponds to a voice of the particular user. In some embodiments, the speech-based service is initiated upon determining that the sound input includes predetermined content and that the sound input corresponds to a voice of the particular user. In some embodiments, the speech-based service is initiated in the restricted access mode upon determining that the sound input includes predetermined content and that the sound input does not correspond to a voice of the particular user. In some embodiments, the method includes outputting an audio prompt including the name of the particular user upon determining that the sound input corresponds to a voice of the particular user.

[0020] In some embodiments, determining whether the sound input includes the predetermined content includes comparing a representation of the sound input to a reference representation, and determining that the sound input includes the predetermined content if the representation of the sound input matches the reference representation. In some embodiments, a match is determined if the representation of the sound input matches the reference representation with a predetermined confidence value. In some embodiments, the method includes receiving a plurality of sound inputs including the sound input, and iteratively adjusting the reference representation using each one of the plurality of sound inputs in response to determining that each sound input includes the predetermined content.

[0021] In some implementations, the method includes determining whether the electronic device is in a predetermined orientation and, upon determining that the electronic device is in the predetermined orientation, enabling a predetermined mode of the voice trigger. In some implementations, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing down, and the predetermined mode is a standby mode. In some implementations, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing up, and the predetermined mode is a listening mode.

[0022] Some embodiments provide a method of operating a voice trigger, the method being performed on an electronic device including one or more processors and a memory storing instructions executed by the one or more processors. The method includes operating the voice trigger in a first mode. The method further includes determining if the electronic device is within a substantially enclosed space by detecting that one or more of a microphone and a camera of the electronic device are occluded. The method further includes switching the voice trigger to a second mode upon determining that the electronic device is within the substantially enclosed space. In some embodiments, the second mode is a standby mode.

[0023] Some embodiments provide a method of operating a voice trigger. The method is performed in an electronic device that includes one or more processors and a memory that stores instructions executed by the one or more processors. The method includes determining whether the electronic device is in a predetermined orientation, and enabling a predetermined mode of the voice trigger upon determining that the electronic device is in the predetermined orientation. In some embodiments, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing down, and the predetermined mode is a standby mode. In some embodiments, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing up, and the predetermined mode is a listening mode.

[0024] In some embodiments, the electronic device includes a receiving unit configured to receive an audio input and a processing unit coupled to the receiving unit. The processing unit is configured to determine whether at least a portion of the audio input corresponds to a predetermined type of sound, determine whether the audio input includes a predetermined content upon determining that at least a portion of the audio input corresponds to the predetermined type, and initiate a speech-based service upon determining that the audio input includes the predetermined content. In some embodiments, the processing unit is further configured to determine whether the audio input satisfies a predetermined condition before determining whether the audio input corresponds to the predetermined type of sound. In some embodiments, the processing unit is further configured to determine whether the audio input corresponds to a voice of a particular user.

[0025] In some embodiments, the electronic device includes a voice trigger unit configured to operate a voice trigger in a first mode of a plurality of modes and a processing unit coupled to the voice trigger unit. In some embodiments, the processing unit is configured to determine whether the electronic device is in a substantially enclosed space by detecting that one or more of a microphone and a camera of the electronic device are occluded, and switch the voice trigger to a second mode upon determining that the electronic device is in the substantially enclosed space. In some embodiments, the processing unit is configured to determine whether the electronic device is in a predetermined orientation, and enable the predetermined mode of the voice trigger upon determining that the electronic device is in the predetermined orientation.

[0026] According to some embodiments, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided that stores one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for performing any of the methods described herein.

[0027] According to some embodiments, there is provided an electronic device (eg, a portable electronic device) that includes means for performing any of the methods described herein.

[0028] According to some embodiments, there is provided an electronic device (eg, a portable electronic device) that includes a processing unit configured to perform any of the methods described herein.

[0029] According to some embodiments, there is provided an electronic device (e.g., a portable electronic device) that includes one or more processors and memory that stores one or more programs executed by the one or more processors, the one or more programs including instructions for performing any of the methods described herein.

[0030] According to some embodiments there is provided an information processing device for use in an electronic device, the information processing device including means for performing any of the methods described herein. [Brief description of the drawings]

[0031] [Figure 1] FIG. 1 is a block diagram illustrating an environment in which a digital assistant operates, according to some embodiments.

[0032] [Diagram 2] FIG. 1 is a block diagram illustrating a digital assistant client system according to some embodiments.

[0033] [Figure 3A] FIG. 1 is a block diagram illustrating a standalone digital assistant system or a digital assistant server system according to some embodiments.

[0034] [Figure 3B] FIG. 3B is a block diagram illustrating the functionality of the digital assistant shown in FIG. 3A in accordance with some embodiments.

[0035] [Figure 3C] FIG. 2 is a network diagram illustrating a portion of an ontology according to some embodiments.

[0036] [Figure 4] FIG. 1 is a block diagram illustrating components of a voice trigger system according to some embodiments.

[0037] [Diagram 5] 1 is a flow chart illustrating a method for operating a voice trigger system according to some embodiments. [Figure 6] 1 is a flow chart illustrating a method for operating a voice trigger system according to some embodiments. [Figure 7] 1 is a flow chart illustrating a method for operating a voice trigger system according to some embodiments.

[0038] [Figure 8] FIG. 1 is a functional block diagram of an electronic device in accordance with some embodiments. [Figure 9] FIG. 1 is a functional block diagram of an electronic device in accordance with some embodiments.

[0039] Like reference numbers refer to corresponding parts throughout the drawings. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0040] FIG. 1 is a block diagram of a digital assistant operating environment 100 according to some embodiments. The terms “digital assistant”, “virtual assistant”, “intelligent automated assistant”, “voice-based digital assistant”, or “automated digital assistant” refer to any information processing system that interprets natural language input in spoken and / or text form to infer user intent (e.g., identify a task type corresponding to the natural language input) and perform an action based on the inferred user intent (e.g., perform a task corresponding to the identified task type). For example, to act based on the inferred user intent, the system can perform one or more of the following: identify a task flow having steps and parameters designed to fulfill the inferred user intent (e.g., identify a task type), input specific requests from the inferred user intent into the task flow, execute the task flow by invoking a program, method, service, API, or the like (e.g., send a request to a service provider), and generate an output response to the user in an audible (e.g., speech) and / or visual form.

[0041] Specifically, once initiated, the digital assistant system can accept user requests, at least in part, in the form of natural language commands, requests, statements, predications, and / or queries. Generally, a user request seeks either an informational answer or the performance of a task by the digital assistant system. Generally, a satisfactory response to a user request will be either the provision of the requested information answer, the performance of the requested task, or a combination of the two. For example, a user may ask the digital assistant system a question such as "Where am I right now?" Based on the user's current location, the digital assistant may respond with "You are near the West Gate in Central Park." The user may also request the performance of a task by stating, for example, "I want you to invite my friends to my girlfriend's birthday party next week." In response, the digital assistant may confirm the request by generating a voice output of "Yes, right away," and then send appropriate calendar invitations from the user's email address to each of the user's friends listed in the user's electronic address book or contact list. There are many other ways to interact with a digital assistant to request information or the performance of various tasks. In addition to providing verbal responses and taking programmed actions, the digital assistant can also provide other visual or audio forms of responses (e.g., as text, alerts, music, videos, animations, etc.).

[0042] As shown in FIG. 1, in some embodiments, the digital assistant system is implemented according to a client-server model. The digital assistant system includes a client-side portion (e.g., 102a and 102b) (hereinafter, "digital assistant (DA) client 102") that runs on a user device (e.g., 104a and 104b) and a server-side portion 106 (hereinafter, "digital assistant (DA) server 106") that runs on a server system 108. The DA client 102 communicates with the DA server 106 through one or more networks 110. The DA client 102 provides client-side functionality, such as user-responsive input and output processing and communication with the DA server 106. The DA server 106 provides server-side functionality for any number of DA clients 102, each resident on a respective user device 104 (also referred to as a client device or electronic device).

[0043] In some implementations, the DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, an I / O interface 118 to external services, a photo and tag database 130, and a photo tag module 132. The client-facing I / O interface facilitates client-facing input and output processing for the digital assistant server 106. The one or more processing modules 114 utilize the data and models 116 to determine user intent based on natural language input and perform tasks based on the inferred user intent. The photo and tag database 130 stores digital photo fingerprints and, optionally, the digital photos themselves, as well as tags associated with the digital photos. The photo tag module 132 creates and stores tags in association with the photos and / or fingerprints, automatically tags photos, and links tags to locations within the photos.

[0044] In some implementations, the DA server 106 communicates with external services 120 (e.g., navigation service(s) 122-1, messaging service(s) 122-2, information service(s) 122-3, calendar service 122-4, phone service 122-5, photo service(s) 122-6, etc.) through network(s) 110 to complete a task or obtain information. An I / O interface 118 to external services facilitates such communication.

[0045] Examples of user equipment 104 include, but are not limited to, a handheld computer, a wireless personal digital assistant (PDA), a tablet computer, a laptop computer, a desktop computer, a cellular telephone, a smart phone, an enhanced general packet radio service (EGPRS) mobile phone, a media player, a navigation device, a game console, a television, a remote control, or a combination of any two or more of these data processing devices, or any other suitable data processing device. Further details regarding user equipment 104 are provided with respect to the exemplary user equipment 104 shown in FIG.

[0046] Examples of communication network(s) 110 include local area networks (LANs) and wide area networks (WANs), such as the Internet. Communication network(s) 110 may be implemented using any known network protocol, including various wired or wireless protocols, such as Ethernet, Universal Serial Bus (USB), FIREWIRE, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wi-Fi, voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.

[0047] The server system 108 may be implemented on at least one data processing device and / or a distributed network of computers. In some implementations, the server system 108 also utilizes the services of various virtual machines and / or third party service providers (e.g., third party cloud service providers) to provide the underlying computing and / or infrastructure resources of the server system 108.

[0048] The digital assistant system shown in FIG. 1 includes both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), but in some embodiments, the digital assistant system refers only to the server-side portion (e.g., DA server 106). In some embodiments, the functionality of the digital assistant can be implemented as a standalone application installed on a user device. In addition, the distribution of functionality between the client and server portions of the digital assistant may vary depending on the embodiment. For example, in some embodiments, the DA client 102 is a thin client that provides only user-facing input and output processing functionality and leaves all other functions of the digital assistant to the DA server 106. For example, in some embodiments, the DA client 102 is configured to perform or assist one or more functions of the DA server 106.

[0049] 2 is a block diagram of user equipment 104 according to some embodiments. User equipment 104 includes a memory interface 202, one or more processors 204, and a peripherals interface 206. The various components within user equipment 104 are coupled by one or more communication buses or signal lines. User equipment 104 includes various sensors, subsystems, and peripherals that are coupled to the peripherals interface 206. The sensors, subsystems, and peripherals collect information and / or facilitate various functions of user equipment 104.

[0050] For example, in some implementations, a motion sensor 210 (e.g., an accelerometer), a light sensor 212, a GPS receiver 213, a temperature sensor, and a proximity sensor 214 are coupled to the peripherals interface 206 to facilitate orientation, light, and proximity sensing functions. In some implementations, other sensors 216, such as biosensors, barometers, etc., are connected to the peripherals interface 206 to facilitate related functions.

[0051] In some implementations, the user equipment 104 includes a camera subsystem 220 coupled to the peripherals interface 206. In some implementations, an optical sensor 222 of the camera subsystem 220 facilitates camera functions such as taking pictures and recording video clips. In some implementations, the user equipment 104 includes one or more wired and / or wireless communication subsystems 224 that provide communication functions. The communication subsystem 224 typically includes various communication ports, radio frequency receivers and transmitters, and / or optical (e.g., infrared) receivers and transmitters. In some implementations, the user equipment 104 includes an audio subsystem 226 coupled to one or more speakers 228 and one or more microphones 230 to facilitate voice-enabled functions such as voice recognition, voice response, digital recording, and telephone functions. In some implementations, the audio subsystem 226 is coupled to a voice trigger system 400. In some implementations, the voice trigger system 400 and / or audio subsystem 226 includes low power audio circuitry and / or programs (i.e., including hardware and / or software) for receiving and / or analyzing audio input, including, for example, one or more analog-to-digital converters, digital signal processors (DSPs), sound detectors, memory buffers, codecs, etc. In some implementations, the low power audio circuitry (alone or in addition to other components of the user equipment 104) provides voice (or sound) trigger functionality for one or more aspects of the user equipment 104, such as a voice-based digital assistant or other speech-based service. In some implementations, the low power audio circuitry provides voice trigger functionality even when other components of the user equipment 104, such as the processor(s) 204, I / O subsystem 240, memory 250, etc., are turned off and / or in standby mode. This voice trigger system 400 is described in more detail with respect to FIG. 4.

[0052] In some implementations, the I / O subsystem 240 is also coupled to the peripherals interface 206. In some implementations, the user device 104 includes a touchscreen 246 and the I / O subsystem 240 includes a touchscreen controller 242 coupled to the touchscreen 246. If the user device 104 includes the touchscreen 246 and the touchscreen controller 242, the touchscreen 246 and the touchscreen controller 242 are typically configured to detect contact and movement or interruption thereof using any of a number of touch sensing technologies, such as, for example, capacitive, resistive, infrared, surface ultrasonic technology, proximity sensor arrays, and the like. In some implementations, the user device 104 includes a display that does not include a touch-sensitive surface. In some implementations, the user device 104 includes a separate touch-sensitive surface. In some implementations, the user device 104 includes other input controller(s) 244. If the user device 104 includes other input controller(s) 244, the other input controller(s) 244 are typically coupled to other input / control devices 248, such as one or more buttons, rocker switches, thumb wheels, infrared ports, USB ports, and / or pointer devices such as styluses.

[0053] The memory interface 202 is coupled to a memory 250. In some implementations, the memory 250 includes a persistent computer-readable medium, such as high-speed random access memory and / or non-volatile memory (e.g., one or more magnetic disk storage devices, one or more flash memory devices, one or more optical storage devices, and / or other non-volatile solid-state storage devices). In some implementations, the memory 250 stores an operating system 252, a communication module 254, a graphical user interface module 256, a sensor processing module 258, a telephony module 260, and applications 262, and a subset or superset thereof. The operating system 252 includes instructions for handling basic system services and instructions for performing hardware-dependent tasks. The communication module 254 facilitates communication with one or more additional devices, one or more computers, and / or one or more servers. The graphical user interface module 256 facilitates graphic user interface processing. The sensor processing module 258 facilitates sensor-related processing and functions (e.g., processing voice input received using one or more microphones 228). The telephony module 260 facilitates telephony-related processes and functions. The application module 262 facilitates various functions of a user application, such as electronic messaging, web browsing, media processing, navigation, imaging, and / or other processes and functions. In some implementations, the user equipment 104 stores in memory 250 one or more software applications 270-1 and 270-2, each associated with at least one of the external service providers.

[0054] As described above, in some implementations, memory 250 also stores client-side digital assistant instructions (e.g., in digital assistant client module 264) and various user data 266 (e.g., user-specific vocabulary data, preference data, and / or other data, such as the user's electronic address book or contact list, to-do list, shopping list, etc.) to provide the client-side functionality of the digital assistant.

[0055] In some implementations, the digital assistant client module 264 can accept voice input, text input, touch input, and / or gesture input through various user interfaces (e.g., I / O subsystem 244) of the user equipment 104. The digital assistant client module 264 can also provide output in audio, visual, and / or haptic forms. For example, the output can be provided as voice, sound, alerts, text messages, menus, graphics, videos, animations, vibrations, and / or combinations of two or more of the above. In operation, the digital assistant client module 264 communicates with a digital assistant server (e.g., digital assistant server 106, FIG. 1) using the communication subsystem 224.

[0056] In some implementations, the digital assistant client module 264 uses various sensors, subsystems, and peripherals to collect additional information from the surrounding environment of the user device 104 to establish a context associated with the user input. In some implementations, the digital assistant client module 264 provides the context information, or a subset thereof, along with the user input to the digital assistant server (e.g., digital assistant server 106, FIG. 1) to help infer the user's intent.

[0057] In some embodiments, context information that may accompany a user input includes sensor information, such as ambient lighting, ambient noise, ambient temperature, images or video, etc. In some embodiments, the context information also includes the physical state of the device, such as device orientation, device location, device temperature, power levels, speed, acceleration, movement patterns, cellular signal strength, etc. In some embodiments, information related to the software state of the user device 106, such as the user device 104's running processes, installed programs, past and current network activity, background services, error logs, resource usage, etc., is also provided to the digital assistant server (e.g., digital assistant server 106, FIG. 1) as context information associated with the user input.

[0058] In some implementations, the DA client module 264 selectively provides information stored on the user device 104 (e.g., at least a portion of the user data 266) in response to a request from the digital assistant server. In some implementations, the digital assistant client module 264 also elicits additional input from the user via a natural language dialog or other user interface in response to a request by the digital assistant server 106 (FIG. 1). The digital assistant client module 264 passes the additional input to the digital assistant server 106 to assist the digital assistant server 106 in estimating and / or achieving the user intent expressed in the user request.

[0059] In some implementations, memory 250 may include additional or fewer instructions. Additionally, various functions of user equipment 104 may be implemented in hardware and / or firmware, including in the form of one or more signal processing and / or application specific integrated circuits, and thus user equipment 104 need not include all of the modules and applications shown in FIG.

[0060] FIG. 3A is a block diagram of an exemplary digital assistant system 300 (also referred to as a digital assistant) according to some embodiments. In some embodiments, the digital assistant system 300 is implemented on a stand-alone computer system. In some embodiments, the digital assistant system 300 is distributed across multiple computers. In some embodiments, some of the modules and functions of the digital assistant are split into a server portion and a client portion. The client portion resides on a user device (e.g., user device 104) and communicates with the server portion (e.g., server system 108) through one or more networks, e.g., as shown in FIG. 1. In some embodiments, the digital assistant system 300 is an embodiment of the server system 108 (and / or digital assistant server 106) shown in FIG. 1. In some embodiments, the digital assistant system 300 is implemented within a user device (e.g., user device 104, FIG. 1), thereby eliminating the need for a client-server system. It should be noted that digital assistant system 300 is merely one example of a digital assistant system, and that digital assistant system 300 may have more or fewer components than shown, may combine two or more components, or may have a different configuration or arrangement of components. The various components shown in FIG. 3A may be implemented in hardware, software, firmware, or a combination thereof, including one or more signal processing and / or application specific integrated circuits.

[0061] The digital assistant system 300 includes a memory 302, one or more processors 304, an input / output (I / O) interface 306, and a network communication interface 308. These components communicate with each other through one or more communication buses or signal lines 310.

[0062] In some implementations, memory 302 includes a persistent computer-readable medium, such as high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., one or more magnetic disk storage devices, one or more flash memory devices, one or more optical storage devices, and / or other non-volatile solid-state memory devices).

[0063] The I / O interface 306 couples the input / output devices 316 of the digital assistant system 300, such as a display, keyboard, touch screen, and microphone, to the user interface module 322. The I / O interface 306 cooperates with the user interface module 322 to receive user inputs (e.g., voice input, keyboard input, touch input, etc.) and process them accordingly. In some implementations, if the digital assistant is implemented on a standalone user device, the digital assistant system 300 includes any of the components described with respect to the user device 104 in FIG. 2 as well as I / O and communication interfaces (e.g., one or more microphones 230). In some implementations, the digital assistant system 300 represents the server portion of a digital assistant implementation and interacts with the user through a client-side portion that resides on the user device (e.g., the user device 104 shown in FIG. 2).

[0064] In some implementations, the network communication interface 308 includes wired communication port(s) 312 and / or wireless transceiver circuitry 314. The wired communication port(s) receive and transmit communication signals via one or more wired interfaces, such as Ethernet, Universal Serial Bus (USB), FIREWIRE, etc. The wireless circuitry 314 receives and transmits RF and / or optical signals, typically to and from communication networks and other communication devices. The wireless communication may use any of a number of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. The network communication interface 308 enables communication between the digital assistant system 300 and networks such as the Internet, an intranet, and / or wireless networks such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs), and other devices.

[0065] In some implementations, the persistent computer-readable storage medium of the memory 302 stores programs, modules, instructions, and data structures, including all or a subset of an operating system 318, a communications module 320, a user interface module 322, one or more applications 324, and a digital assistant module 326. The one or more processors 304 execute these programs, modules, instructions, and read / write from / to the data structures.

[0066] Operating system 318 (e.g., Darwin®, RTXC®, LINUX®, UNIX®, OS X®, iOS®, Windows®, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage control, power management, etc.) and facilitating communication between the various hardware, firmware, and software components.

[0067] The communications module 320 facilitates communications between the digital assistant system 300 and other devices through the network communications interface 308. For example, the communications module 320 can communicate with the communications module 254 of the device 104 shown in FIG. 2. The communications module 320 also includes various software components for processing data received by the wireless circuitry 314 and / or the wired communications port 312.

[0068] In some implementations, the user interface module 322 receives commands and / or input from a user via the I / O interface 306 (e.g., from a keyboard, touch screen, and / or microphone) and provides user interface objects on a display.

[0069] Applications 324 include programs and / or modules configured to be executed by one or more processors 304. For example, if the digital assistant system is implemented on a standalone user device, applications 324 may include user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 300 is implemented on a server farm, applications 324 may include, for example, resource management applications, diagnostic applications, or scheduling applications.

[0070] The memory 302 also stores a digital assistant module (i.e., the server portion of the digital assistant) 326. In some implementations, the digital assistant module 326 includes the following submodules, or a subset or superset thereof: an input / output processing module 328, a speech-to-text (STT) processing module 330, a natural language processing module 332, a dialog flow processing module 334, a task flow processing module 336, a service processing module 338, and a photo module 132. Each of these processing modules has access to one or more of the following data and models of the digital assistant 326, or a subset or superset thereof: an ontology 360, a vocabulary index 344, user data 348, a classification module 349, a disambiguation module 350, a task flow model 354, a service model 356, a photo tagging module 358, a search module 360, and a local tag / photo storage 362.

[0071] In some implementations, using the processing modules (e.g., input / output processing module 328, STT processing module 330, natural language processing module 332, dialog flow processing module 334, task flow processing module 336, and / or service processing module 338), data, and models implemented in the digital assistant module 326, the digital assistant system 300 does at least some of the following: Identify the user's intent expressed in the natural language input received from the user, actively elicit and obtain the information necessary to fully infer the user's intent (e.g., by disambiguating words, names, intents, etc.), determine a task flow to satisfy the inferred intent, and execute the task flow to satisfy the inferred intent. In some implementations, the digital assistant also takes appropriate action when a satisfactory response is not or cannot be provided to the user for various reasons.

[0072] In some implementations, as described below, the digital assistant system 300 processes the natural language input to identify the user's intent from the natural language input to tag the digital photos and tag the digital photos with appropriate information. In some implementations, the digital assistant system 300 also performs other photo-related tasks, such as searching for digital photos using the natural language input, automatically tagging photos, etc. As shown in FIG. 3B, in some implementations, the I / O processing module 328 interacts with the user through the I / O device 316 of FIG. 3A or with the user equipment (e.g., the user equipment 104 of FIG. 1) through the network communication interface 308 of FIG. 3A to obtain user input (e.g., speech input) and provide a response to the user input. The I / O processing module 328 optionally obtains context information associated with the user input from the user equipment along with or immediately after receiving the user input. The context information includes user-specific data, vocabulary, and / or preferences associated with the user input. In some implementations, the context information also includes information about the software and hardware state of the device (e.g., user device 104 in FIG. 1) at the time the user request is received and / or the user's surrounding environment at the time the user request is received. In some implementations, the I / O processing module 328 also sends follow-up questions to the user about the user request and receives answers from the user. In some implementations, when a user request is received by the I / O processing module 328 and the user request includes voice input, the I / O processing module 328 forwards the voice input to a speech-to-text (STT) processing module 330 for speech-to-text conversion.

[0073] In some implementations, the speech-to-text processing module 330 receives speech input (e.g., a user's utterance captured in an audio recording) through the I / O processing module 328. In some implementations, the speech-to-text processing module 330 uses various acoustic and language models to recognize the speech input as a sequence of phonemes and ultimately as a sequence of words or tokens written in one or more languages. The speech-to-text processing module 330 is implemented using any suitable speech recognition technique, acoustic model, and language model, such as hidden Markov models, dynamic time warping (DTW)-based speech recognition, and other statistical and / or analytical techniques. In some implementations, the speech-to-text processing may be performed at least in part by a third-party service or on the user's device. Once the speech-to-text processing module 330 obtains the results of the speech-to-text processing (e.g., a sequence of words or tokens), it passes the results to the natural language processing module 332 for intent estimation. The natural language processing module 332 ("natural language processor") of the digital assistant 326 takes the string of words or tokens ("token string") generated by the speech-to-text processing module 330 and attempts to associate the token string with one or more "actionable intents" recognized by the digital assistant. As used herein, an "actionable intent" refers to a task that can be executed by the digital assistant 326 and / or the digital assistant system 300 (FIG. 3A) and has an associated task flow implemented in the task flow model 354. The associated task flow is a sequence of programmed actions and steps that the digital assistant system 300 takes to execute a task. The scope of the digital assistant system's capabilities depends on the number and type of task flows implemented and stored in the task flow model 354, or in other words, the number and type of "actionable intents" that the digital assistant system 300 recognizes. However, the effectiveness of the digital assistant system 300 also depends on the digital assistant system's ability to infer the correct "actionable intent(s)" from a user request expressed in natural language.

[0074] In some implementations, in addition to the string of words or tokens obtained from the speech-to-text processing module 330, the natural language processor 332 also receives contextual information associated with the user request (e.g., from the I / O processing module 328). The natural language processor 332 optionally uses the contextual information to clarify, complement, and / or further clarify information contained within the token string received from the speech-to-text processing module 330. Contextual information includes, for example, user preferences, hardware and / or software state of the user equipment, sensor information collected before, during, or immediately after the user request, previous interactions (e.g., dialogue) between the digital assistant and the user, and the like.

[0075] In some embodiments, the natural language processing is based on an ontology 360. The ontology 360 is a hierarchical structure that includes multiple nodes, each of which represents either an "actionable intention" or an "attribute" that is related to one or more of a "multiple actionable intentions" or other "multiple attributes." As described above, an "actionable intention" represents a task that the digital assistant system 300 has the ability to perform (e.g., a task that is "actionable" or can be the subject of performance). An "attribute" represents a parameter associated with an actionable intention or a sub-aspect of another attribute. A link between an actionable intention node and an attribute node in the ontology 360 defines how a parameter represented by the attribute node is related to a task represented by the actionable intention node. In some embodiments, the ontology 360 is composed of actionable intention nodes and attribute nodes. In the ontology 360, each actionable intention node is linked to one or more attribute nodes directly or through one or more intermediate attribute nodes. Similarly, each attribute node is linked to one or more actionable intention nodes directly or through one or more intermediate attribute nodes. For example, the ontology 360 shown in FIG. 3C includes an actionable intent node, “Restaurant Reservation” node. The attribute nodes “Restaurant”, “Date / Time” (for reservation), and “Number of Parties” are each directly connected to the “Restaurant Reservation” node (i.e., the actionable intent node). In addition, the attribute nodes “Cuisine”, “Price Range”, “Phone Number” and “Location” are subnodes of the attribute node “Restaurant”, and each are connected to the “Restaurant Reservation” node via the intermediate attribute node “Restaurant”. For another example, the ontology 360 shown in FIG. 3C also includes another actionable intent node, “Set Reminder” node. The attribute nodes “Date / Time” (for reminder setting) and “Theme” (for reminder) are each connected to the “Set Reminder” node. Because the attribute node “Date / Time” is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node “Date / Time” is connected to both the “Restaurant Reservation” node and the “Set Reminder” node in the ontology 360.

[0076] An actionable intent node, together with its connected concept nodes, can be described as a "domain." In this description, each domain is associated with a respective actionable intent and refers to the set of nodes (and relationships between them) associated with a particular actionable intent. For example, ontology 360 shown in FIG. 3C includes an example of a restaurant reservation domain 362 and an example of a reminder domain 364 within ontology 360. The restaurant reservation domain includes an actionable intent node "restaurant reservation," attribute nodes "restaurant," "date / time," and "number of parties," and sub-attribute nodes "cuisine," "price range," "phone number," and "location." The reminder domain 364 includes an actionable intent node "reminder settings," and attribute nodes "theme" and "date / time." In some implementations, ontology 360 is composed of multiple domains. Each domain can share one or more attribute nodes with one or more other domains. For example, the attribute node for “date / time” can be associated with many other domains (e.g., scheduling domain, travel booking domain, movie ticket domain, etc.) in addition to the restaurant reservation domain 362 and reminder domain 364. Although FIG. 3C shows two example domains within the ontology 360, the ontology 360 may include other domains (i.e., actionable intents) such as “initiate a call,” “get directions,” “schedule a meeting,” “send a message,” and “provide an answer to a question,” “tag a photo,” etc. For example, the domain for “send a message” is associated with the actionable intent node for “send a message,” and may further include attribute nodes such as “recipient(s),” “message type,” and “message body.” The attribute node “recipient” may be further defined by sub-attribute nodes such as, for example, “recipient name” and “message address.”

[0077] In some implementations, ontology 360 includes all domains (and therefore actionable intents) that the digital assistant can understand and act upon. In some implementations, ontology 360 may be modified, such as by adding or removing domains or nodes, or by changing relationships between nodes in ontology 360.

[0078] In some implementations, nodes associated with multiple related actionable intents may be clustered under a “super domain” in the ontology 360. For example, a “travel” super domain may include a cluster of travel-related attribute nodes and actionable intent nodes. The travel-related actionable intent nodes may include “book a flight”, “book a hotel”, “rent a car”, “get directions”, “find sights”, etc. Actionable intent nodes under the same super domain (e.g., the “travel” super domain) may share many attribute nodes. For example, the actionable intent nodes for “book a flight”, “book a hotel”, “rent a car”, “get directions”, and “find sights” may share one or more of the attribute nodes “departure location”, “destination”, “departure date / time”, “arrival date / time”, and “number of people involved”.

[0079] In some implementations, each node in the ontology 360 is associated with a set of words and / or phrases related to the attribute or actionable intent represented by that node. The respective set of words and / or phrases associated with each node is the so-called "vocabulary" associated with that node. The respective set of words and / or phrases associated with each node can be stored in the vocabulary index 344 (FIG. 3B) in association with the attribute or actionable intent represented by that node. For example, returning to FIG. 3B, the vocabulary associated with a node for the attribute "restaurant" may include words such as "food," "drink," "dish," "hungry," "eat," "pizza," "fast food," "meal," etc. As another example, the vocabulary associated with a node for the actionable intent "initiate a phone call" may include words and phrases such as "call," "phone," "dial," "ring," "call this number," "make a call to," etc. The vocabulary index 344 optionally includes words and phrases in different languages. In some implementations, the natural language processor 332 shown in FIG. 3B receives a token sequence (e.g., a text string) from the speech-to-text processing module 330 and determines which nodes are implied by the words in the token sequence. In some implementations, if a word or phrase in the token sequence is found to be associated with one or more nodes in the ontology 360 (via the vocabulary index 344), the word or phrase will "trigger" or "activate" those nodes. If multiple nodes are "triggered," the natural language processor 332 will select one of the possible intents as the task (or type of task) that the user intends the digital assistant to perform, based on the amount and / or relative importance of the activated nodes. In some implementations, the domain with the most "triggered" nodes is selected.In some embodiments, the domain with the highest confidence value (e.g., based on the relative importance of its various triggered nodes) is selected. In some embodiments, the domain is selected based on a combination of the number and importance of triggered nodes. In some embodiments, additional factors are also considered when selecting a node, such as whether the digital assistant system 300 has previously correctly interpreted a similar request from the user.

[0080] In some implementations, the digital assistant system 300 also stores the names of specific entities in the vocabulary index 344. Therefore, when one of these names is found in a user request, the natural language processor 332 can recognize that the name refers to a specific instance of an attribute or subattribute in the ontology. In some implementations, the names of specific entities are names of businesses, restaurants, people, movies, and the like. In some implementations, the digital assistant system 300 can search and identify specific entity names from other data sources, such as the user's address book, contact list, movie database, musician database, and / or restaurant database. In some implementations, when the natural language processor 332 identifies a word in a token string as the name of a specific entity (such as a name in the user's address book or contact list), the word is given additional weight in selecting an actionable intent in the ontology for the user request. For example, if the word "Mr. Santo" is recognized from a user request, and the last name "Santo" is found in vocabulary index 344 as one of the contacts in the user's contact list, then the user request likely corresponds to the "send a message" or "initiate a call" domain. As another example, if the word "ABC Cafe" is found in a user request, and the term "ABC Cafe" is found in vocabulary index 344 as the name of a particular restaurant in the user's city, then the user request likely corresponds to the "restaurant reservation" domain.

[0081] User data 348 includes user-specific information, such as user-specific vocabulary, user preferences, user addresses, the user's default and secondary languages, the user's contact list, and other short-term or long-term information about each user. The natural language processor 332 can use the user-specific information to supplement information contained in the user input to further clarify the user's intent. For example, in response to a user request "invite my friends to my birthday party," the natural language processor 332 can access user data 348 to determine who the "friends" are and when and where the "birthday party" will be held, instead of requiring the user to explicitly provide such information in the user's request.

[0082] In some implementations, the natural language processor 332 includes a classification module 349. In some implementations, the classification module 349 determines whether each of one or more terms in a text string (e.g., corresponding to a voice input associated with a digital photograph) is one of an entity, an action, or a location, as described in more detail below. In some implementations, the classification module 349 classifies each term of the one or more terms as being one of an entity, an action, or a location. Once the natural language processor 332 identifies an actionable intent (or domain) based on the user request, the natural language processor 332 generates a structured query to represent the identified actionable intent. In some implementations, the structured query includes parameters for one or more nodes in the domain of the actionable intent, with at least some of the parameters populated with specific information and requirements specified in the user request. For example, a user may say, "Please make a dinner reservation for me at a sushi restaurant at 7 o'clock." In this case, the natural language processor 332 may be able to accurately identify the actionable intent as "restaurant reservation" based on the user input. According to the ontology, a structured query for the “restaurant reservation” domain may include parameters such as {cuisine}, {time}, {date}, {number of parties}, and the like. Based on the information contained in the user's utterance, the natural language processor 332 may generate a partial structured query for the restaurant reservation domain. Here, the partial structured query includes the parameters {cuisine="sushi"} and {time="7pm"}. However, in this example, the user's utterance does not include enough information to complete a structured query associated with the domain. Therefore, other required parameters such as {number of parties} and {date} are not specified in the structured query based on the currently available information. In some implementations, the natural language processor 332 adds the received context information to some parameters of the structured query.For example, if a user requests sushi restaurants "near me," the natural language processor 332 may add GPS coordinates from the user device 104 to the {location} parameter in the structured query.

[0083] In some implementations, the natural language processor 332 passes the structured query (including any completed parameters) to a task flow processing module 336 ("task flow processor"). The task flow processor 336 is configured to receive the structured query from the natural language processor 332, complete the structured query, and perform actions required to "complete" the user's final request. In some implementations, the various steps required to complete these tasks are provided in the task flow model 354. In some implementations, the task flow model 354 includes steps for obtaining additional information from the user and task flows for performing actions associated with the actionable intent. As described above, to complete the structured query, the task flow processor 336 may need to initiate additional dialogue with the user to obtain additional information and / or disambiguate potentially ambiguous utterances. If such dialogue is required, the task flow processor 336 invokes the dialog processing module 334 (dialog processor) to engage in a dialogue with the user. In some implementations, the dialog processing module 334 determines how (and / or when) to ask the user for additional information and receives and processes user responses. In some implementations, the dialog processing module 334 provides questions to the user and receives answers from the user through the I / O processing module 328. For example, the dialog processing module 334 presents dialog output to the user via audio and / or visual output and receives input from the user via verbal or physical (e.g., touch gesture) responses. Continuing with the example above, when the task flow processor 336 invokes the dialog processor 334 to determine "number of parties" and "date" information for a structured query associated with the domain "restaurant reservation," the dialog processor 334 generates questions such as "for how many people?" and "on what date?" to pass to the user.Upon receiving an answer from the user, the dialog processing module 334 passes the information to the task flow processor 336 to add or complete the missing information in the structured query.

[0084] In some cases, the task flow processor 336 may receive a structured query with one or more ambiguous attributes. For example, a structured query for the "send message" domain may indicate that the intended recipient is "Bob" and the user may have multiple contacts named "Bob." The task flow processor 336 will request that the dialog processor 334 disambiguate this attribute of the structured query. As a result, the dialog processor 334 may ask the user, "Which Bob?" and display (or read out) a list of contacts named "Bob" from which the user can choose.

[0085] In some implementations, the dialog processor 334 includes a disambiguation module 350. In some implementations, the disambiguation module 350 disambiguates one or more ambiguous terms (e.g., one or more ambiguous terms of a text string corresponding to a voice input associated with a digital photograph). In some implementations, the disambiguation module 350 determines whether a first term of the one or more terms has multiple possible meanings, prompts a user for additional information about the first term, receives the additional information from the user in response to the prompt, and identifies an entity, action, or location associated with the first term in response to the additional information.

[0086] In some implementations, the disambiguation module 350 disambiguates pronouns. In such implementations, the disambiguation module 350 identifies one of the one or more terms as a pronoun and determines the noun to which the pronoun refers. In some implementations, the disambiguation module 350 determines the noun to which the pronoun refers using a contact list associated with the user of the electronic device. Alternatively, or in addition, the disambiguation module 350 determines the noun to which the pronoun refers as the name of an entity, action, or place identified in a previous voice input associated with the previously tagged digital photograph. Alternatively, or in addition, the disambiguation module 350 determines the noun to which the pronoun refers as the name of a person identified based on a previous voice input associated with the previously tagged digital photograph. In some implementations, the disambiguation module 350 accesses information obtained from one or more sensors (e.g., the proximity sensor 214, the light sensor 212, the GPS receiver 213, the temperature sensor 215, the motion sensor 210) of the handheld electronic device (e.g., the user device 104) to determine the meaning of one or more of the terms. In some implementations, the disambiguation module 350 identifies two terms each associated with either an entity, an action, or a location. For example, a first of the two terms refers to a person and a second of the two terms refers to a location. In some implementations, the disambiguation module 350 identifies three terms each associated with either an entity, an action, or a location.

[0087] Once the task flow processor 336 completes the structured query for the actionable intent, the task flow processor 336 proceeds to execute the final task associated with the actionable intent. In response, the task flow processor 336 executes steps and instructions in the task flow model according to the specific parameters contained in the structured query. For example, a task flow model for an actionable intent of "restaurant reservation" may include steps and instructions for contacting a restaurant and actually requesting a reservation for a specific number of parties at a specific time. For example, using a structured query such as {restaurant reservation, restaurant=ABC Cafe, date=3 / 12 / 2012, time=7pm, number of parties=5}, the task flow processor 336 may (1) log into a server at ABC Cafe or into a restaurant reservation system configured to accept reservations for multiple restaurants such as ABC Cafe, (2) enter date, time, and number of parties information into a form on the website, (3) submit the form, and (4) calendar the reservation into the user's calendar. In another example, described in more detail below, the task flow processor 336 may, for example, cooperate with the photo module 132 to execute steps and instructions associated with tagging or searching digital photos in response to voice input. In some implementations, the task flow processor 336 employs the assistance of a service processing module 338 ("service processor") to complete a task requested in the user input or provide an answer to information requested in the user input. For example, the service processor 338 may perform actions on behalf of the task flow processor 336, such as making a phone call, setting a calendar entry, invoking a map search, invoking or interacting with other user applications installed on the user equipment, and invoking or interacting with third party services (e.g., a restaurant reservation portal, a social networking website or service, a banking portal, etc.).In some embodiments, the protocols and application programming interfaces (APIs) required by each service may be specified by a respective service model in service models 356. Service processor 338 accesses the appropriate service model for the service and generates a request for the service according to the protocols and APIs required by the service associated with the service model.

[0088] For example, if a restaurant enables an online reservation service, the restaurant may present a service model that specifies the parameters required to make a reservation and an API for communicating the values ​​of the required parameters to the online reservation service. When requested by the task flow processor 336, the service processor 338 may establish a network connection with the online reservation service using the web address stored in the service model 356 and transmit the required reservation parameters (e.g., time, date, number of parties) to the online reservation interface in a format conforming to the online reservation service's API.

[0089] In some implementations, the natural language processor 332, the dialog processor 334, and the task flow processor 336 are used collaboratively and iteratively to infer and clarify the user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response that accomplishes the user's intent (e.g., providing an output to the user or completing a task).

[0090] In some embodiments, after all tasks necessary to fulfill the user's request have been performed, the digital assistant 326 formulates an acknowledgment response and sends the response back to the user through the I / O processing module 328. If the user request asks for an informational answer, the acknowledgment response presents the requested information to the user. In some embodiments, the digital assistant also requests the user to indicate whether the user is satisfied with the response created by the digital assistant 326.

[0091] Attention is now directed to FIG. 4, a block diagram illustrating components of a voice trigger system 400, according to some embodiments. (The voice trigger system 400 is not limited to speech, and the embodiments described herein apply equally to non-voice sounds.) The voice trigger system 400 is comprised of various components, modules, and / or software programs within the electronic device 104. In some embodiments, the voice trigger system 400 includes a noise detector 402, a sound type detector 404, a trigger sound detector 406, a speech-based service 408, and an audio subsystem 226, each coupled to an audio bus 401. In some embodiments, more or fewer modules are used. The sound detectors 402, 404, and 406 may be referred to as modules, and may include hardware (e.g., circuits, memory, processors, etc.), software (e.g., programs, software on a chip, firmware, etc.), and / or any combination thereof, to perform the functions described herein. In some implementations, the sound detectors are communicatively, programmatically, physically, and / or operatively coupled to one another (e.g., via a communications bus), as shown by the dashed lines in Figure 4. (For ease of illustration, Figure 4 shows each sound detector coupled only to adjacent sound detectors. It will be understood that each sound detector may be coupled to any of the other sound detectors as well.)

[0092] In some implementations, the audio subsystem 226 includes a codec 410, an audio digital signal processor (DSP) 412, and a memory buffer 414. In some implementations, the audio subsystem 226 is coupled to one or more microphones 230 (FIG. 2) and one or more speakers 228 (FIG. 2). The audio subsystem 226 provides sound input to sound detectors 402, 404, 406 and speech-based services 408 (as well as other components or modules, such as a telephone and / or a telephone baseband subsystem) for processing and / or analysis. In some implementations, the audio subsystem 226 is coupled to an external audio system 416 including at least one microphone 418 and at least one speaker 420.

[0093] In some implementations, the speech-based service 408 is a voice-based digital assistant and corresponds to one or more components or functions of the digital assistant system described above in connection with FIGS. 1-3C. In some implementations, the speech-based service is a speech-to-text service, a dictation service, etc., and in some implementations, the noise detector 402 monitors an audio channel to determine whether the sound input from the audio subsystem 226 meets a predetermined condition, such as an amplitude threshold. An audio channel corresponds to a stream of audio information received by one or more sound-capturing devices, such as one or more microphones 230 (FIG. 2). An audio channel may refer to audio information regardless of its processing state, or may refer to a particular hardware that is processing and / or transmitting the audio information. For example, an audio channel may refer to analog electrical impulses from the microphone 230 (and / or the circuitry through which they are propagated), as well as a digitally encoded audio stream resulting from the processing of the analog electrical impulses (e.g., by the audio subsystem 226 and / or any other audio processing system of the electronic device 104).

[0094] In some implementations, the predetermined condition is whether the sound input exceeds a certain volume at a predetermined time. In some implementations, the noise detector uses a time domain analysis of the sound input, which requires relatively less computational and battery resources compared to other types of analysis (e.g., as performed by sound type detector 404, trigger word detector 406, and / or speech-based service 408). In some implementations, other types of signal processing and / or audio analysis are used, including, for example, frequency domain analysis. When the noise detector 402 determines that the sound input meets the predetermined condition, it activates an upstream sound detector, such as sound type detector 404 (e.g., by providing a control signal to initiate one or more processing routines and / or by providing power to the upstream sound detector). In some implementations, the upstream sound detector is activated in response to other conditions being met. For example, in some implementations, the upstream sound detector is activated in response to determining that the device is not stored in an enclosed space (e.g., based on a light detector detecting a threshold level of light).

[0095] The sound type detector 404 monitors the audio channel to determine whether the sound input corresponds to a particular type of sound, such as a sound characteristic of a human voice, a whistle, a clap, etc. The type of sound that the sound type detector 404 is configured to recognize corresponds to the particular trigger sound(s) that the voice trigger is configured to recognize. In embodiments where the trigger sound is a spoken word or phrase, the sound type detector 404 includes a "voice activity detector" (VAD). In some embodiments, the sound type detector 404 uses a frequency domain analysis of the sound input. For example, the sound type detector 404 generates a spectrogram of the received sound input (e.g., using a Fourier transform) to analyze the spectral components of the sound input and determine whether the sound input appears to correspond to a particular type or classification of sound (e.g., a human voice). Thus, in embodiments where the trigger sound is a spoken word or phrase, if the audio channel picks up background sounds (e.g., traffic noise) rather than human voice, the VAD will not activate the trigger sound detector 406. In some implementations, sound type detector 404 remains enabled as long as a predetermined condition of any downstream sound detector (e.g., noise detector 402) is satisfied. For example, in some implementations, sound type detector 404 remains enabled as long as the sound input includes sound above a predetermined amplitude threshold (as determined by noise detector 402) and is disabled when the sound falls below the predetermined threshold. In some implementations, once activated, sound type detector 404 remains enabled until a condition is satisfied, such as expiration of a timer (e.g., 1, 2, 5, or 10 seconds, or any other suitable duration), expiration of a particular number of on / off cycles of sound type detector 404, or occurrence of an event (e.g., the amplitude of the sound falls below a second threshold, as determined by noise detector 402 and / or sound type detector 404).

[0096] As described above, when the sound type detector 404 determines that the sound input corresponds to a predetermined type of sound, it activates an upstream sound detector, such as the trigger sound detector 406 (e.g., by providing a control signal to initiate one or more processing routines and / or by providing power to the upstream sound detector).

[0097] The trigger sound detector 406 is configured to determine whether the sound input includes at least a portion of a particular predefined content (e.g., at least a portion of a trigger word, phrase, or sound). In some implementations, the trigger sound detector 406 compares a representation of the sound input (the "input representation") to one or more reference representations of the trigger word. If the input representation matches at least one of the one or more reference representations with an acceptable confidence value, the trigger sound detector 406 initiates the speech-based service 408 (e.g., by providing a control signal to initiate one or more processing routines and / or by providing power to an upstream sound detector). In some implementations, the input representation and the one or more reference representations are spectrograms (or a mathematical representation thereof), which represent how the spectral density of a signal changes over time. In some implementations, the representations are other types of audio signatures or voiceprints. In some implementations, initiating the speech-based service 408 includes bringing one or more circuits, programs, and / or processors out of standby mode and invoking the sound-based service. The sound-based service then prepares to provide more comprehensive speech recognition, speech-to-text processing, and / or natural language processing. In some implementations, the voice trigger system 400 includes voice recognition functionality so that it can determine whether the sound input corresponds to the voice of a particular person, such as the owner / user of the device. For example, in some implementations, the sound type detector 404 uses voice printing techniques to determine that the sound input was spoken by an authorized user. Voice recognition and voice printing are described in more detail in commonly owned U.S. patent application Ser. No. 13 / 053,144, which is incorporated herein by reference in its entirety. In some implementations, voice recognition is included in any of the sound detectors described herein (e.g., the noise detector 402, the sound type detector 404, the trigger sound detector 406, and / or the speech-based service 408).In some implementations, voice authentication is implemented as a separate module from the sound detectors described above (e.g., as voice authentication module 428, FIG. 4 ) and may be operatively located after noise detector 402, after sound type detector 404, after trigger sound detector 406, or in any other suitable location.

[0098] In some implementations, trigger sound detector 406 remains enabled as long as the conditions of any downstream sound detector(s) (e.g., noise detector 402 and / or sound type detector 404) are satisfied. For example, in some implementations, trigger sound detector 406 remains enabled as long as the sound input includes a sound above a predetermined threshold (as detected by noise detector 402). In some implementations, it remains enabled as long as the sound input includes a particular type of sound (as detected by sound type detector 404). In some implementations, it remains enabled as long as both of the aforementioned conditions are satisfied.

[0099] In some implementations, once activated, the trigger sound detector 406 remains enabled until a condition is met, such as the expiration of a timer (e.g., 1, 2, 5, or 10 seconds, or any other suitable duration), the expiration of a particular number of on / off cycles of the trigger sound detector 406, or the occurrence of an event (e.g., the amplitude of the sound drops below a second threshold). In some implementations, when one sound detector activates another, both sound detectors remain enabled. However, sound detectors may be enabled or disabled multiple times, and it is not necessary for all of the downstream (e.g., low power and / or sophisticated) sound detectors to be enabled (or for each condition to be met) in order for an upstream sound detector to be enabled. For example, in some implementations, after the noise detector 402 and the sound type detector 404 determine that each of their conditions have been met and the trigger sound detector 406 has been activated, one or both of the noise detector 402 and the sound type detector 404 are disabled and / or in standby mode while the trigger sound detector 406 is operating. In other embodiments, both noise detector 402 and sound type detector 404 (or one or the other) remain enabled during operation of trigger sound detector 406. In various embodiments, different combinations of sound detectors are enabled at different times, and whether one is enabled or disabled may depend on the state of the other sound detectors or may be independent of the state of the other sound detectors.

[0100] 4 illustrates three separate sound detectors, each configured to detect a different type of sound input, and various implementations of the voice trigger may use more or fewer sound detectors. For example, in some implementations, only trigger sound detector 406 is used. In some implementations, trigger sound detector 406 is used in conjunction with either noise detector 402 or sound type detector 404. In some implementations, all of detectors 402-406 are used. In some implementations, additional sound detectors are included as well.

[0101] Moreover, different combinations of sound detectors may be used at different times. For example, the particular combination of sound detectors and how they interact may depend on one or more conditions, such as the context or operating state of the device. As one specific example, when the device is plugged in (and thus not solely dependent on battery power), the trigger sound detector 406 is enabled while the noise detector 402 and sound type detector 404 are kept disabled. As another example, when the device is in a pocket or backpack, all sound detectors are disabled. By cascading sound detectors as described above, where detectors requiring more power are invoked only when needed by detectors requiring less power, a power-saving voice trigger function may be provided. Further power savings are achieved by operating one or more of the sound detectors according to a duty cycle, as described above. For example, in some implementations, the noise detector 402 operates according to a duty cycle to effectively provide continuous noise detection even when the noise detector is at least temporarily disabled. In some implementations, the noise detector 402 is on for 10 milliseconds and off for 90 milliseconds. In some implementations, the noise detector 402 is on for 20 ms and off for 500 ms, although other on and off durations are possible.

[0102] In some implementations, once the noise detector 402 detects noise during its "on" interval, the noise detector 402 remains on and further processes and / or analyzes the sound input. For example, the noise detector 402 may be configured to activate an upstream sound detector upon detecting a sound above a predetermined amplitude for a predetermined period of time (e.g., 100 milliseconds). Thus, once the noise detector 402 detects a sound above a predetermined amplitude during its 10 millisecond "on" interval, it does not immediately enter an "off" interval. Instead, the noise detector 402 remains enabled and continues to process the sound input and determine whether a threshold is exceeded for a predetermined total duration (e.g., 100 milliseconds).

[0103] In some implementations, sound type detector 404 operates according to a duty cycle. In some implementations, sound type detector 404 is on for 20 milliseconds and off for 100 milliseconds. Other on and off durations are possible. In some implementations, sound type detector 404 can determine whether a sound input corresponds to a predefined type of sound during the "on" interval of its duty cycle. Thus, if sound type detector 404 determines during its "on" interval that a sound is of a particular type, sound type detector 404 activates trigger sound detector 406 (or any other upstream sound detector). Alternatively, in some implementations, if sound type detector 404 detects a sound that can correspond to a predefined type during its "on" interval, the detector does not immediately enter an "off" interval. Instead, sound type detector 404 remains enabled and continues to process sound input and determine whether it corresponds to a predefined type of sound. In some implementations, once the sound detector determines that a predetermined type of sound has been detected, it activates the trigger sound detector 406 to further process the sound input and determine if a trigger sound has been detected. Like the noise detector 402 and the sound type detector 404, in some implementations, the trigger sound detector 406 operates according to a duty cycle. In some implementations, the trigger sound detector 406 is on for 50 milliseconds and off for 50 milliseconds. Other on and off durations are possible. If the trigger sound detector 406 detects during its "on" interval that there is a sound that can correspond to a trigger sound, the detector does not immediately enter an "off" interval. Instead, the trigger sound detector 406 remains enabled and continues to process the sound input and determine if it includes a trigger sound. In some implementations, once such a sound is detected, the trigger sound detector 406 remains enabled and processes the audio for a predetermined duration, such as 1, 2, 5, or 10 seconds, or any other suitable duration. In some implementations, the duration is selected based on the length of the particular trigger word or sound it is configured to detect. For example, if the trigger phrase is "To SIRI," the trigger word detector will operate for approximately two seconds to determine if the sound input contains the phrase.

[0104] In some implementations, some of the sound detectors are operated according to a duty cycle, while others operate continuously when enabled. For example, in some implementations, only the first sound detector is operated according to a duty cycle (e.g., noise detector 402 of FIG. 4), and the upstream sound detectors are operated continuously once activated. In some other implementations, noise detector 402 and sound type detector 404 are operated according to a duty cycle, while trigger sound detector 406 is operated continuously. Whether a particular sound detector is operated continuously or according to a duty cycle depends on one or more conditions, such as the context or operating state of the device. In some implementations, all of the sound detectors operate continuously once activated if the device is connected to a power source and does not rely solely on battery power. In other implementations, noise detector 402 (or any of the sound detectors) operates according to a duty cycle when the device is in a pocket or backpack (e.g., as determined by a sensor and / or microphone signal), but operates continuously if it is determined that the device may not be stored. In some implementations, whether a particular sound detector is operated continuously or according to a duty cycle depends on the battery charge level of the device. For example, the noise detector 402 operates continuously when the battery charge is above 50% and according to a duty cycle when the battery charge is below 50%. In some implementations, the voice trigger includes noise, echo, and / or sound cancellation functions (collectively referred to as noise cancellation). In some implementations, noise cancellation is performed by the audio subsystem 226 (e.g., by the audio DSP 412). Noise cancellation reduces or removes unwanted noise or sounds from the sound input before it is processed by the sound detector. In some implementations, the unwanted noise is background noise from the user's environment, such as a fan or clicking from a keyboard. In some implementations, the unwanted noise is any sound above or below a certain amplitude or frequency.For example, in some implementations, sounds above the typical human speech range (e.g., 3,000 Hz) are filtered out or removed from the signal. In some implementations, multiple microphones (e.g., microphone 230) are used to help determine which components of the received sound should be reduced and / or removed. For example, in some implementations, audio subsystem 226 uses beamforming techniques to identify sounds or portions of the sound input that originate from a single point in space (e.g., the user's mouth). Audio subsystem 226 then focuses on sounds that are received equally by all microphones (e.g., background sounds that are not originating from any particular direction) by removing them from the sound input.

[0105] In some implementations, the DSP 412 is configured to cancel or remove from the sound input the sound being output by the device on which the digital assistant is operating. For example, if the audio subsystem 226 is outputting music, radio, podcasts, voice output, or any other audio content (e.g., via the speaker 228), the DSP 412 removes any of the output sound picked up by the microphone and included in the sound input. Thus, the sound input does not include this output sound (or at least includes less of the output sound). Accordingly, the sound input provided to the sound detector is cleaner and more accurate trigger. Aspects of noise cancellation are described in more detail in commonly assigned U.S. Pat. No. 7,272,224, which is incorporated herein by reference in its entirety.

[0106] In some implementations, different sound detectors require the sound input to be filtered and / or pre-processed in different ways. For example, in some implementations, the noise detector 402 is configured to analyze time-domain sound signals between 60 and 20,000 Hz, and the sound type detector is configured to perform frequency-domain analysis of sound between 60 and 3,000 Hz. Thus, in some implementations, the audio DSP 412 (and / or other audio DSPs of the device 104) pre-processes the received sound according to the respective needs of the sound detector. In some implementations, the sound detectors, on the other hand, are configured to filter and / or pre-process the sound from the audio subsystem 226 according to their specific needs. In such cases, the audio DSP 412 may still perform noise cancellation before providing the sound input to the sound detector. In some implementations, the context of the electronic device is used to help determine if and how a voice trigger is activated. For example, if the device is in a pocket, purse, or backpack, the user is unlikely to invoke a speech-based service, such as a voice-based digital assistant. Also, a user is unlikely to invoke a speech-based service during a loud rock concert. A user is unlikely to invoke a speech-based service at a particular time (e.g., late at night). However, there are contexts in which a user may well invoke a speech-based service using a voice trigger. For example, a user may well use a voice trigger while driving, when alone, at work, etc. Various techniques are used to determine the device's context. In various implementations, the device uses information from any one or more of the following components or sources to determine the device's context: GPS receiver, light sensor, microphone, proximity sensor, orientation sensor, inertial sensor, camera, communication circuitry and / or antenna, charging circuitry and / or power circuitry, switch position, temperature sensor, compass, accelerometer, calendar, user preferences, etc.The device context can then be used to adjust whether and how the voice trigger operates. For example, in certain contexts, the voice trigger is disabled (or operated in a different mode) as long as the context is maintained. For example, in some implementations, the voice trigger is disabled when the phone is in a certain orientation (e.g., face down on a surface), during a certain period of time (e.g., between 10:00 PM and 8:00 AM), when the phone is in "silent" or "do not disturb" mode (e.g., based on switch position, mode setting, or user preference), when the device is in a substantially enclosed space (e.g., a pocket, bag, purse, drawer, or glove box), when the device is near other devices that have voice triggers and / or speech-based services (e.g., based on proximity sensors, voice / wireless / infrared communications), etc. In some implementations, instead of being disabled, the voice trigger system 400 is operated in a low power mode (e.g., by operating the noise detector 402 according to a duty cycle with an "on" interval of 10 milliseconds and an "off" interval of 5 seconds). In some implementations, the audio channel is monitored less frequently when the voice trigger system 400 is operated in a low power mode. In some implementations, the voice trigger uses a different sound detector or combination of sound detectors when in a low power mode than when in a normal mode. (The voice trigger may enable many different modes or operating states, each of which may use different amounts of power, and different implementations use them according to their particular designs.)

[0107] On the other hand, if the device is in some other context, the voice trigger will be active (or operated in a different mode) as long as the context is maintained. For example, in some implementations, the voice trigger will remain active while the phone is connected to a power source, while the phone is in a certain orientation (e.g., facing up on a surface), during a certain period of time (e.g., between 8:00AM and 10:00PM), while the device is moving and / or in a vehicle (e.g., based on a GPS signal, a BLUETOOTH connection, or while connected to a vehicle, etc.). Aspects of detecting confirmation that the device is in a vehicle are described in more detail in commonly assigned U.S. Provisional Patent Application No. 61 / 657,744, which is incorporated herein by reference in its entirety. Various examples of methods for determining a particular context are provided below. In various embodiments, these and other contexts are detected using different techniques and / or sources of information.

[0108] As discussed above, whether the voice trigger system 400 is enabled (e.g., listening) may depend on the physical orientation of the device. In some implementations, the voice trigger is enabled when the device is placed on a surface "face up" (e.g., with the display and / or touchscreen surface visible) and / or disabled when placed on a surface "face down". This provides a user with an easy way to enable and / or disable the voice trigger without having to navigate through settings menus, switches, or buttons. In some implementations, the device detects whether it is placed face up or face down on a surface using a light sensor (e.g., based on the difference in incident light on the front and back of the device 104), a proximity sensor, a magnetic sensor, an accelerometer, a gyroscope, a tilt sensor, a camera, etc. In some implementations, other operating modes, settings, parameters, or preferences are affected by the orientation and / or position of the device. In some implementations, the particular trigger sound, word, or phrase that the voice trigger is listening for depends on the orientation and / or position of the device. For example, in some implementations, the voice trigger listens for a first trigger word, phrase, or sound when the device is in one orientation (e.g., face up on a surface) and listens for a different trigger word, phrase, or sound when the device is in another orientation (e.g., face down). In some implementations, the trigger phrase for the face down orientation is longer and / or more complex than that for the face up orientation. Thus, a user can place the device face down when other people are around or in a noisy environment and still activate the voice trigger while also reducing fraudulent acknowledgements that would occur more frequently for shorter or simpler trigger words. As one specific example, the face up trigger phrase may be "To SIRI," while the face down trigger phrase may be "To SIRI, this is Andrew, please activate." A longer trigger phrase also provides the sound detector and / or voice authenticator with a longer voice sample to process and / or analyze, thus increasing the accuracy of the voice trigger and reducing fraudulent acknowledgements.

[0109] In some implementations, the device 104 detects if the device is in a vehicle (e.g., an automobile). Voice triggers are particularly useful for invoking speech-based services when the user is in a vehicle because they help reduce the physical interaction required to operate the device and / or speech-based services. Indeed, one advantage of a voice-based digital assistant is that it can be used to perform tasks when it is impossible or dangerous to see and touch the device. Thus, voice triggers may be used when the device is in a vehicle so that the user does not need to touch the device to invoke the digital assistant. In some implementations, the device determines that it is in the vehicle by detecting that it is connected to and / or paired with the vehicle, such as through BLUETOOTH communication (or other wireless communication) or a docking connector or cable. In some implementations, the device determines that it is in the vehicle by determining the location and / or velocity of the device (e.g., using a GPS receiver, accelerometer, and / or gyroscope). For example, it may be determined that the device is likely to be inside a vehicle because it is traveling above 20 miles per hour and is determined to be moving along a road, and the voice trigger may continue to remain enabled and / or in a high power or high sensitivity state.

[0110] In some implementations, the device detects whether the device is stored (e.g., in a pocket, purse, bag, drawer, etc.) by determining whether it is in a substantially enclosed space. In some implementations, the device uses a light sensor (e.g., a dedicated ambient light sensor and / or a camera) to determine whether it is stored. For example, in some implementations, the device is likely stored when the light sensor detects low or no light. In some implementations, the time of day and / or the location of the device are also taken into consideration. For example, when high light levels are expected (e.g., during the day), if the light sensor detects low light levels, the device is stored and the voice trigger system 400 may not be needed. Thus, the voice trigger system 400 goes into a low power or standby state. In some implementations, the difference in light detected by sensors located on opposite sides of the device may be used to determine its location and therefore whether it is stored. Specifically, if the device is not stored in a pocket or bag but is placed on a table or surface, the user may attempt to activate the voice trigger. When the device is placed face down (or face up) on a surface such as a table or desk, one side of the device is covered and receives little or no light while the other surface is exposed to ambient light. Thus, if the light sensors on the front and back of the device detect significantly different light levels, the device is determined to be not stored. On the other hand, if the light sensors on the opposing sides detect the same or similar light levels, the device is determined to be stored in a substantially enclosed space. Also, if both light sensors detect low light levels during the day (or if the device expects the phone to be in a bright environment), the device is determined to be stored with a high degree of confidence.

[0111] In some implementations, other techniques are used (instead of or in addition to optical sensors) to determine if the device is stored. For example, in some implementations, the device emits one or more sounds (e.g., tones, clicks, pings, etc.) from a speaker or transducer (e.g., speaker 228) and monitors one or more microphones or transducers (e.g., microphone 230) to detect echoes of the omitted sound(s). (In some implementations, the device emits inaudible signals, such as sounds outside the range of human hearing.) From the echoes, the device determines characteristics of the surrounding environment. For example, a relatively large environment (e.g., inside a room or car) reflects sound differently than a relatively small, enclosed environment (e.g., a pocket, purse, bag, drawer, etc.).

[0112] In some implementations, the voice trigger system 400 is operated differently when it is close to other devices (such as other devices with voice triggers and / or speech-based services) than when it is far from the other devices. This can be useful, for example, to shut off or desensitize the voice trigger system 400 when many devices are close to each other, so that when one person speaks a trigger word, other surrounding devices are not triggered as well. In some implementations, the device determines its proximity to other devices using RFID, proximity communication, infrared / acoustic signals, etc. As mentioned above, voice triggers are particularly useful when the device is operated in a hands-free mode, such as when the user is driving. In such cases, users often use external audio systems, such as wired or wireless headsets, watches with speakers and / or microphones, vehicle built-in microphones and speakers, etc., to make calls or dictate text input without having to hold the device close to their face. For example, wireless headsets and vehicle audio systems may connect to the electronic device using BLUETOOTH® communication, or any other suitable wireless communication. However, it can be inefficient for a voice trigger to monitor audio received via a wireless audio accessory because of the power required to maintain an open audio channel with the wireless accessory. In particular, a wireless headset can hold enough power in its battery to provide several hours of continuous talk time, and is therefore well-suited to conserve the battery for when the headset is needed for actual communication, instead of using it to simply monitor ambient audio and wait for potential trigger sounds. Moreover, a wired external headset accessory may require excessive power compared to an on-board microphone alone, and keeping the headset's microphone active drains the device's battery charge power. This is especially true considering that the ambient audio received by a wireless or wired headset usually consists mostly of silence or irrelevant sounds.Thus, in some implementations, the voice trigger system 400 monitors audio from the on-device microphone 230 even if the device is coupled to an external microphone (wired or wireless). If the voice trigger then detects a trigger word, the device initiates an active audio link with the external microphone and receives subsequent sound input (such as a command to a voice-based digital assistant) via the external microphone rather than the on-device microphone 230. If certain conditions are met, a valid communication link may be maintained between the external audio system 416 (which may be communicatively coupled to the device 104 via wire or wireless) and the device, and the voice trigger system 400 may listen for the trigger sound via the external audio system 416 instead of (or in addition to) the on-device microphone 230. For example, in some implementations, the movement characteristics of the electronic device and / or the external audio system 416 (e.g., as determined by an accelerometer, gyroscope, etc. on each device) are used to determine whether the voice trigger system 400 should monitor background sounds using the microphone 230 on the device or the external microphone 418. Specifically, the difference in movement between the device and the external audio system 416 provides information about whether the external audio system 416 is actually in use. For example, if both the device and the wireless headset are moving (or not moving) substantially equally, it may be determined that the headset is not in use or not being worn. This may occur, for example, because both devices are close to each other and are idle (e.g., resting on a table or in a pocket, bag, purse, drawer, etc.). Accordingly, under these conditions, the voice trigger system 400 monitors the microphone on the device because it is unlikely that the headset is actually in use. If there is a difference in movement between the wireless headset and the device, it is determined that the user is wearing the headset.These conditions may occur, for example, while the headset is worn on the user's head (where at least a small amount of movement would be possible even if the wearer is relatively stationary) because the device is placed (e.g., on a surface or in a bag). Under these conditions, the headset is considered to be worn, so the voice trigger system 400 maintains a valid communication link and monitors the headset's microphone 418 instead of (or in addition to) the microphone 230 on the device. Because this technique looks at differences in the movement of the device and the headset, any movement common to both devices is cancelled out. This may be useful, for example, when a user is using a headset in a moving vehicle where the device (e.g., a mobile phone) is in a cup holder, on an empty seat, or in the user's pocket, and the headset is worn on the user's head. Once any movement common to both devices is cancelled out (e.g., vehicle movement), the relative movement (if any) of the headset compared to the device may be determined to determine whether the headset is likely in use (or whether the headset is not being worn). Although the above description refers to wireless headsets, similar techniques apply to wired headsets as well.

[0113] Because human voices vary widely, it may be necessary or useful to tune a voice trigger to improve its accuracy in recognizing a particular user's voice. Also, a person's voice may change over time due to, for example, natural voice changes due to illness, aging, or hormonal changes. Thus, in some implementations, the voice trigger system 400 can adapt its voice and / or sound recognition profile to a particular user or group of users. As described above, a sound detector (e.g., sound type detector 404 and / or trigger sound detector 406) may be configured to compare a representation of a sound input (e.g., a sound or utterance provided by a user) to one or more reference representations. For example, if the input representation matches the reference representations at a predetermined confidence level, the sound detector determines that the sound input corresponds to a predetermined type of sound (e.g., sound type detector 404) or that the sound input includes predetermined content (e.g., trigger sound detector 406). To tune the voice trigger system 400, in some embodiments, the device calibrates a reference representation to which the input representation is compared. In some embodiments, the reference representation is calibrated (or created) as part of a voice enrollment or "training" procedure, where the user outputs a trigger sound several times to allow the device to tune (or create) the reference representation. The device then creates the reference representation using the person's actual voice.

[0114] In some implementations, the device uses trigger sounds received under normal use conditions to adjust the reference representation. After a successful voice triggering event (e.g., a sound input is found that meets all of the triggering criteria), for example, the device uses information from the sound input to adjust and / or tune the reference representation. In some implementations, only sound inputs that are determined to meet all or part of the triggering criteria with a certain confidence level are used to adjust the reference representation. Thus, if the voice trigger has low confidence that a sound input corresponds to or contains a trigger sound, that voice input may be ignored for purposes of adjusting the reference representation. On the other hand, in some implementations, sound inputs that meet the voice trigger system 400 with a low confidence level are used to adjust the reference representation.

[0115] In some implementations, the device 104 iteratively adjusts the reference representation (using these or other techniques) as more and more sound input is received to accommodate for slight changes in the user's voice over time. For example, in some implementations, the device 104 (and / or associated devices or services) adjusts the reference representation after each successful triggering event. In some implementations, the device 104 analyzes the sound input associated with each successful triggering event, determines whether the reference representation should be adjusted based on that input (e.g., if certain conditions are met), and adjusts the reference representation only if it is appropriate to do so. In some implementations, the device 104 maintains a running average of the reference representation over time. In some implementations, the voice trigger system 400 detects sounds that do not meet one or more of the triggering criteria (e.g., as determined by one or more of the sound detectors), which may be an actual attempt to do so by a legitimate user. For example, the voice trigger system 400 may be configured to respond to a trigger phrase such as "to SIRI," but if the user's voice changes (e.g., due to illness, aging, change in accent / tone, etc.), the voice trigger system 400 may not recognize the user's attempt to activate the device. (This may also occur if the voice trigger system 400 is set to a default condition and / or the voice trigger system 400 has not been properly tuned to the user's particular voice, such as if the user has not initialized or gone through a training procedure to customize the voice trigger system 400 for the user's voice.) If the voice trigger system 400 does not respond to the user's first attempt to activate the voice trigger, the user will likely repeat the trigger phrase. The device will detect that these repeated sound inputs are similar to each other and / or to the trigger phrase (even if they are not similar enough to cause the voice trigger system 400 to activate the speech-based service).If such conditions are met, the device determines that the sound inputs correspond to a legitimate attempt to activate the voice trigger system 400. In response, in some implementations, the voice trigger system 400 uses those received sound inputs to adjust one or more aspects of the voice trigger system 400 so that similar utterances by the user are recognized as legitimate triggers in the future. In some implementations, these sound inputs are used to adapt the voice trigger system 400 only if a certain condition or combination of conditions is met. For example, in some implementations, the sound inputs are used to adapt the voice trigger system 400 when a predetermined number of sound inputs are received consecutively (e.g., 2, 3, 4, 5, or any other suitable number), when the sound inputs are sufficiently similar to a reference representation, when the sound inputs are sufficiently similar to each other, when the sound inputs are close to each other (e.g., received within a predetermined time period and / or at or near a predetermined interval), and / or under any combination of these or other conditions. In some cases, the voice trigger system 400 may detect one or more sound inputs that do not meet one or more of the triggering criteria followed by a manual initiation (e.g., by pressing a button or icon) of a speech-based service. In some implementations, the voice trigger system 400 determines that the sound inputs do in fact correspond to a failed voice triggering attempt because the speech-based service was initiated shortly after receiving the sound inputs. In response, the voice trigger system 400 uses those received sound inputs to adjust one or more aspects of the voice trigger system 400, as described above, so that utterances by the user are recognized as legitimate triggers in the future.

[0116] While the above accommodation techniques refer to adjusting the reference representation, other aspects of the trigger sound detection techniques may be adjusted in the same or similar manner in addition to or instead of adjusting the reference representation. For example, in some embodiments, the device adjusts how the sound input is filtered and / or what filters are applied to the sound input, such as focusing on and / or reducing a particular frequency or frequency range of the sound input. In some embodiments, the device adjusts the algorithm used to compare the input representation to the reference representation. For example, in some embodiments, one or more terms of a mathematical function used to determine the difference between the input representation and the reference representation are changed, added, or removed, or replaced with a different mathematical function. In some embodiments, accommodation techniques such as those described above require more resources than the voice trigger system 400 can or is configured to provide. In particular, the sound detector may not have the amount or type of, or access to, processor, data, or memory required to perform iterative adaptation of the reference representation and / or sound detection algorithm (or any other suitable aspect of the voice trigger system 400). Thus, in some implementations, one or more of the above adaptation techniques are performed by a more powerful processor, such as an application processor (e.g., processor(s) 204), or by a different device (e.g., server system 108). However, the voice trigger system 400 is designed to operate even when the application processor is in standby mode. Thus, the sound input used to adapt the voice trigger system 400 is received when the application processor is not available and cannot process the sound input. Accordingly, in some implementations, the sound input is stored by the device so that it can be further processed and / or analyzed after receipt. In some implementations, the sound input is stored in a memory buffer 414 of the audio subsystem 226.In some implementations, the sound input is stored in a system memory (e.g., memory 250, FIG. 2) using direct memory access (DMA) techniques (e.g., including using a DMA engine to copy or move data without having to wake up the application processor). The stored sound input is then provided to or accessed by the application processor (or server system 108, or another suitable device) such that upon wake-up, the application processor can perform one or more of the adaptation techniques described above. In some implementations,

[0117] 5-7 are flow diagrams depicting a method for operating a voice trigger according to certain embodiments. The method is optionally governed by instructions stored in a computer memory or persistent computer-readable storage medium (e.g., memory 250 of client device 104, memory 302 associated with digital assistant system 300) and executed by one or more processors of one or more computer systems of the digital assistant system, including but not limited to server system 108 and / or user device 104a. The computer-readable storage medium may include magnetic or optical disk storage devices, solid-state storage devices such as flash memory, or other non-volatile memory device(s). The computer-readable instructions stored on the computer-readable storage medium may include one or more of source code, assembly language code, object code, or other instruction formats interpreted and executed by one or more processors. In various embodiments, some operations of the respective methods shown in each figure may be combined and / or the order of some operations may be changed from the order. Also, in some implementations, operations shown in separate figures and / or described in association with separate methods may be combined to form other methods, and operations described in association with the same figures and / or methods may be separated into different methods. Moreover, in some implementations, one or more operations in a method are performed by modules of the digital assistant system 300 and / or electronic device (e.g., user device 104), including, for example, the natural language processing module 332, the dialog flow processing module 334, the audio subsystem 226, the noise detector 402, the sound type detector 404, the trigger sound detector 406, the speech-based service 408, and / or any submodules thereof. FIG. 5 illustrates a method 500 for operating a voice trigger system (e.g., the voice trigger system 400 of FIG. 4, FIG. 4) according to some implementations. In some implementations, the method 500 is performed in an electronic device including one or more processors and a memory that stores instructions executed by the one or more processors (e.g., electronic device 104).The electronics receives (502) audio input, which may correspond to speech (e.g., a word, phrase, or sentence), a human pronunciation (e.g., whistling, clicking, finger snapping, clapping, etc.), or any other sound (e.g., electronically generated chirps, mechanical noise makers, etc.). In some implementations, the electronics receives the audio input via audio subsystem 226 (e.g., including codec 410, audio DSP 412, and buffer 414, as well as microphones 230 and 418, described in connection with FIG. 4).

[0118] In some implementations, the electronics determines whether the sound input meets a predefined condition (504). In some implementations, the electronics applies a time domain analysis to the sound input to determine whether the sound input meets a predefined condition. For example, the electronics analyzes the sound input over a period of time to determine whether the sound amplitude reaches a predefined level. In some implementations, the threshold is met if the amplitude (e.g., volume) of the sound input meets and / or exceeds a predefined threshold. In some implementations, it is met if the sound input meets and / or exceeds a predefined threshold for a predefined time. As described in more detail below, in some implementations, determining whether the sound input meets a predefined condition (504) is performed by a third sound detector (e.g., noise detector 402). (The third sound detector is used in this case to distinguish this sound detector from the other sound detectors (e.g., the first sound detector and the second sound detector described below) and does not necessarily indicate any operating position or order of the sound detectors.)

[0119] The electronic device determines whether the sound input corresponds to a predetermined type of sound (506). As discussed above, sounds are categorized into various "types" based on certain identifiable sound characteristics. Determining whether the sound input corresponds to a predetermined type includes determining whether the sound input includes or exhibits characteristics of the particular type. In some implementations, the predetermined type of sound is a human voice. In such implementations, determining whether the sound input corresponds to a human voice includes determining whether the sound input includes a frequency characteristic of a human voice (508). As described in more detail below, in some implementations, determining whether the sound input corresponds to a predetermined type of sound (506) is performed by a first sound detector (e.g., sound type detector 404). Upon determining that the sound input corresponds to a predetermined type of sound, the electronic device determines whether the sound input includes predetermined content (510). In some implementations, the predetermined content corresponds to one or more predetermined phonemes (512). In some implementations, the one or more predetermined phonemes comprise at least one word. In some implementations, the predefined content is a sound (e.g., a whistle, a click, or a clap). In some implementations, determining 510 whether the sound input includes the predefined content is performed by a second sound detector (e.g., trigger sound detector 406), as described below.

[0120] Upon determining that the audio input includes the predetermined content, the electronic device initiates (514) the speech-based service. In some embodiments, the speech-based service is a voice-based digital assistant, as described in more detail above. In some embodiments, the speech-based service is a dictation service, where the speech input is converted to text and included and / or displayed in a text input field (e.g., an email, text message, word processing, or note-taking application, etc.). In embodiments where the speech-based service is a voice-based digital assistant, when the voice-based digital assistant is initiated, a prompt (e.g., a sound or speech prompt) is issued to the user indicating that the user can provide voice input and / or commands to the digital assistant. In some embodiments, initiating the voice-based digital assistant includes enabling an application processor (e.g., processor(s) 204, FIG. 2), initiating one or more programs or modules (e.g., digital assistant client module 264, FIG. 2), and / or establishing a connection to a remote server or device (e.g., digital assistant server 106, FIG. 1).

[0121] In some implementations, the electronic device determines (516) whether the sound input corresponds to the voice of the particular user. For example, one or more voice authentication techniques are applied to the sound input to determine whether it corresponds to the voice of an authorized user of the device. Voice authentication techniques are described in detail above. In some implementations, the voice authentication is performed by one of the sound detectors (e.g., trigger sound detector 406). In some implementations, the voice authentication is performed by a dedicated voice authentication module (including any suitable hardware and / or software). In some implementations, a sound-based service is initiated in response to determining that the sound input includes predetermined content and that the sound input corresponds to the voice of the particular user. Thus, for example, a sound-based service (e.g., a voice-based digital assistant) is initiated only if a trigger word or phrase is spoken by an authorized user. This reduces the likelihood that the service may be invoked by unauthorized users and can be particularly useful when multiple electronic devices are in close proximity, as utterance of a trigger sound by one user does not activate a voice trigger for another user.

[0122] In some implementations where the speech-based service is a voice-based digital assistant, in response to determining that the sound input includes predetermined content but does not correspond to the voice of a particular user, the voice-based digital assistant is initiated in a limited access mode. In some implementations, the limited access mode allows the digital assistant to access only a subset of the data, services, and / or functionality that the digital assistant may otherwise provide. In some implementations, the limited access mode corresponds to a write-only mode (e.g., such that non-regular users of the digital assistant cannot access data from the calendar, task list, contacts, photos, email, text messages, etc.). In some implementations, the limited access mode corresponds to a sandboxed instance of the speech-based service, preventing the speech-based service from reading from or writing to the user's data, such as user data 266 on device 104 (FIG. 2) or any other device (e.g., user data 348 of FIG. 3A, which may be stored on a remote server, such as server system 108 of FIG. 1).

[0123] In some implementations, in response to determining that the sound input includes the predetermined content and that the sound input corresponds to the voice of the particular user, the voice-based digital assistant outputs a prompt that includes the name of the particular user. For example, once the particular user is identified via voice authentication, the voice-based digital assistant may output a prompt such as "Peter, what can I do for you?" instead of a more general prompt such as a tone, beep, or non-proprietary voice prompt. As described above, in some implementations, the first sound detector determines if the sound input corresponds to a predetermined type of sound (at step 506) and the second sound detector determines if the sound detector includes the predetermined content (at step 510). In some implementations, the first sound detector consumes less power in operation than the second sound detector, e.g., because the first sound detector uses less processor-intensive technology than the second sound detector. In some implementations, the first sound detector is sound type detector 404 and the second sound detector is trigger sound detector 406, both of which are described above in connection with FIG. 4. In some implementations, during these operations, the first sound detector and / or the second sound detector periodically monitor the audio channel according to a duty cycle, as described above in connection with FIG.

[0124] In some implementations, the first sound detector and / or the sound detectors perform a frequency domain analysis of the sound input. For example, these sound detectors perform a Laplace transform, a Z transform, or a Fourier transform to generate a frequency spectrum or determine a spectral density of the sound input or a portion thereof. In some implementations, the first sound detector is a voice activity detector configured to determine whether the sound input includes frequencies that are characteristic of a human voice (or other features, aspects, or aspects of the sound input that are characteristic of a human voice).

[0125] In some implementations, the second sound detector is off or disabled until the first sound detector detects a sound input of a predetermined type. Accordingly, in some implementations, the method 500 includes activating the second sound detector in response to determining that the sound input corresponds to the predetermined type. (In other implementations, the second sound detector is activated in response to other conditions or is continuously operated regardless of the determination from the first sound detector.) In some implementations, activating the second sound detector includes enabling hardware and / or software (e.g., including circuits, processors, programs, memory, etc.). In some implementations, the second sound detector is operated (e.g., enabled and monitoring an audio channel) for at least a predetermined time period after activation. For example, when the first sound detector determines that the sound input corresponds to a predetermined type (e.g., includes a human voice), the second sound detector is activated to determine whether the sound input also includes a predetermined content (e.g., a trigger word). In some implementations, the predetermined time period corresponds to a duration of the predetermined content. Thus, if the predetermined content is the phrase "Dear SIRI," the predetermined time will be long enough to determine if the phrase was uttered (e.g., 1 or 2 seconds, or any other suitable duration). If the predetermined content is longer, such as the phrase "Dear SIRI, wake up and help," the predetermined time will be longer (e.g., 5 seconds, or another suitable duration). In some implementations, the second sound detector is activated as long as the first sound detector detects a sound corresponding to the predetermined type. In such implementations, for example, as long as the first sound detector detects a human voice in the sound input, the second sound detector processes the sound input and determines if it includes the predetermined content.

[0126] As discussed above, in some implementations, the third sound detector (e.g., noise detector 402) determines whether the sound input satisfies a predetermined condition (at step 504). In some implementations, the third sound detector consumes less power during operation than the first sound detector. In some implementations, the third sound detector periodically monitors the audio channel according to a duty cycle, as described above with respect to FIG. 4. Also, in some implementations, the third sound detector performs a time domain analysis of the sound input. In some implementations, the third sound detector consumes less power than the first sound detector because the time domain analysis is less processor intensive than the frequency domain analysis applied by the second sound detector.

[0127] Similar to the above discussion regarding activating the second sound detector (e.g., trigger sound detector 406) in response to a determination by the first sound detector (e.g., sound type detector 404), in some implementations the first sound detector is activated in response to a determination by the third sound detector (e.g., noise detector 402). For example, in some implementations the sound type detector 404 is activated in response to a determination by the noise detector 402 that the sound input meets a predetermined condition (e.g., exceeds a particular volume for a sufficient duration). In some implementations, activating the first sound detector includes enabling hardware and / or software (including, e.g., circuits, processors, programs, memory, etc.). In other implementations, the first sound detector is activated in response to other conditions or is continuously operated. In some implementations, the device stores (518) at least a portion of the sound input in memory. In some implementations, the memory is a buffer 414 of the audio subsystem 226 (FIG. 4). The stored sound input allows the device to process the sound input in non-real-time. For example, in some implementations, one or more of the sound detectors read and / or receive the stored sound input and process the stored sound input. This may be particularly useful if an upstream sound detector (e.g., trigger sound detector 406) is not activated until midway through receipt of the sound input by the audio subsystem 226. In some implementations, the stored portion of the sound input is provided (520) to the speech-based service when the speech-based service is initiated. Thus, the speech-based service can copy, process, or otherwise operate on the stored portion of the sound input, even though the speech-based service is not fully operational until the portion of the sound input is received. In some implementations, the stored portion of the sound input is provided to an accommodation module of the electronic device.

[0128] In various embodiments, steps 516-520 occur at different locations within method 500. For example, in some embodiments, one or more of steps 516-520 occur between steps 502 and 504, between steps 510 and 514, or at any other suitable location.

[0129] FIG. 6 illustrates a method 600 for operating a voice trigger system (e.g., voice trigger system 400 of FIG. 4, FIG. 4) according to some embodiments. In some embodiments, method 600 is performed in an electronic device including one or more processors and a memory that stores instructions executed by the one or more processors (e.g., electronic device 104). The electronic device determines whether it is in a predefined orientation (602). In some embodiments, the electronic device detects its orientation using a light sensor (including a camera), a microphone, a proximity sensor, a magnetic sensor, an accelerometer, a gyroscope, a tilt sensor, etc. For example, the electronic device determines whether it is placed face-down or face-up on a surface by comparing the amount or brightness of light incident on a front camera sensor to the amount or brightness of light incident on a rear camera sensor. If the amount and / or brightness detected by the front camera is sufficiently greater than that detected by the rear camera, the electronic device is determined to be face-up. On the other hand, if the amount and / or brightness detected by the rear camera is sufficiently greater than that detected by the front camera, the electronic device is determined to be face-down. Upon determining that the electronic device is in a predetermined orientation, the electronic device enables a predetermined mode of the voice trigger (604). In some implementations, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing down, and the predetermined mode is a standby mode (606). For example, in some implementations, when the smart phone or tablet is placed on a table or desk with the screen facing down, the voice trigger goes into standby mode (e.g., powered off) to prevent unintentional activation of the voice trigger.

[0130] However, in some implementations, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing up, and the predetermined mode is listening mode 608. Thus, for example, if the smart phone or tablet is placed on a table or desk with the screen facing up, the voice trigger will be in listening mode and can respond to the user upon detecting the trigger.

[0131] FIG. 7 illustrates a method 700 for operating a voice trigger (e.g., voice trigger system 400, FIG. 4) according to some embodiments. In some embodiments, the method 700 is performed on an electronic device that includes one or more processors and a memory that stores instructions that are executed by the one or more processors (e.g., electronic device 104). The electronic device operates (702) a voice trigger (e.g., voice trigger system 400) in a first mode. In some embodiments, the first mode is a normal listening mode.

[0132] The electronic device determines 704 whether it is in a substantially enclosed space by detecting that one or more of the electronic device's microphone and camera are occluded. In some implementations, the substantially enclosed space includes a pocket, a purse, a bag, a drawer, a glove box, a briefcase, etc.

[0133] As described above, in some implementations, the device detects that the microphone is occluded by emitting one or more sounds (e.g., tones, clicks, pings, etc.) from a speaker or transducer and monitoring one or more microphones or transducers to detect an echo of the omission sound(s). For example, a relatively large environment (e.g., a room or car) reflects sound differently than a relatively small, substantially enclosed environment (e.g., a purse or pocket). Thus, when the device detects that the microphone (or the speaker that emitted the sound) is occluded based on the echo (or lack of echo), the device determines that it is in a substantially enclosed space. In some implementations, the device detects that the microphone is occluded by detecting that the microphone picks up sounds characteristic of enclosed spaces. For example, if the device is in a pocket, the microphone may detect a characteristic soft noise caused by the microphone touching or being in close proximity to the fabric of the pocket. In some implementations, the device detects that the camera is occluded based on the level of light received by the sensor or by determining whether it can obtain a focused image. For example, if the camera sensor detects a low level of light during a time when high levels of light are expected (e.g., during the day), the device determines that the camera is occluded and that the device is in a substantially enclosed space. As another example, the camera may attempt to obtain a focused image on its sensor. Typically, this is difficult when the camera is in a very dark location (e.g., a pocket or backpack) or is too close to the subject it is trying to focus on (e.g., in a purse or backpack). Thus, if the camera is unable to obtain a focused image, it determines that the device is in a substantially enclosed space.

[0134] Upon determining that the electronic device is within a substantially enclosed space, the electronic device switches the voice trigger to a second mode (706). In some implementations, the second mode is a standby mode (708). In some implementations, when in standby mode, the voice trigger system 400 continues to monitor ambient sounds but does not respond to received sounds regardless of whether the voice trigger system 400 is otherwise activated. In some implementations, in standby mode, the voice trigger system 400 is disabled and does not process sounds to detect trigger sounds. In some implementations, the second mode includes operating one or more sound detectors of the voice trigger system 400 according to a different duty cycle than the first mode. In some implementations, the second mode includes operating a different combination of sound detectors than the first mode.

[0135] In some embodiments, the second mode corresponds to a more sensitive monitoring mode, allowing the voice trigger system 400 to detect and respond to trigger sounds even within the substantially enclosed space. In some embodiments, once the voice trigger has switched to the second mode, the device periodically determines whether the electronic device is still within the substantially enclosed space by detecting whether one or more of the electronic device's microphone and camera are occluded (e.g., using any of the techniques described above with respect to step (704)). If the device is still within the substantially enclosed space, the voice trigger system 400 remains in the second mode. In some embodiments, once the device is removed from the substantially enclosed space, the electronic device switches the voice trigger back to the first mode.

[0136] According to some implementations, FIG. 8 shows a functional block diagram of an electronic device 800 configured according to the principles of the present invention as described above. The functional blocks of the device can be implemented by hardware, software, or a combination of hardware and software to execute the principles of the present invention. It is understood by those skilled in the art that the functional blocks described in FIG. 8 can be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description in this specification may support any possible combination or division, or further definition of the functional blocks described in this specification.

[0137] As shown in FIG. 8 , electronic device 800 includes a sound receiving unit 802 configured to receive a sound input. Electronic device 800 also includes a processing unit 806 coupled to speech receiving unit 802. In some implementations, processing unit 806 includes a noise detector 808, a sound type detector 810, a trigger sound detector 812, a service initiation unit 814, and a voice authentication unit 816. In some implementations, noise detector 808 corresponds to noise detector 402 described above and is configured to perform any of the operations described above for noise detector 402. In some implementations, sound type detector 810 corresponds to sound type detector 404 described above and is configured to perform any of the operations described above for sound type detector 404. In some implementations, trigger sound detector 812 corresponds to trigger sound detector 406 described above and is configured to perform any of the operations described above for trigger sound detector 406. In some implementations, voice authentication unit 816 corresponds to voice authentication module 428 described above and is configured to perform any of the operations described above for voice authentication module 428. The processing unit 806 is configured to determine (e.g., via the sound type detection unit 810) whether at least a portion of the sound input corresponds to a predetermined type of sound, and if it is determined that at least a portion of the sound input corresponds to the predetermined type, determine (e.g., via the trigger sound detection unit 812) whether the sound input includes predetermined content, and if it is determined that the sound input includes the predetermined content, initiate a speech-based service (e.g., via the service initiation unit 814).

[0138] In some implementations, the processing unit 806 is also configured to determine whether the sound input satisfies a predetermined condition (e.g., with the noise detection component 808) before determining whether the sound input corresponds to a predetermined type of sound. In some implementations, the processing unit 806 is also configured to determine whether the sound input corresponds to the voice of a particular user (e.g., with the voice recognition component 816).

[0139] According to some implementations, FIG. 9 shows a functional block diagram of an electronic device 900 configured according to the principles of the present invention as described above. The functional blocks of the device can be implemented by hardware, software, or a combination of hardware and software to execute the principles of the present invention. It is understood by those skilled in the art that the functional blocks described in FIG. 9 can be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description in this specification may support any possible combination or division, or further definition of the functional blocks described in this specification.

[0140] As shown in FIG. 9, the electronic device 900 includes a voice trigger unit 902. The voice trigger unit 902 can be operated in a variety of different modes. In a first mode, the voice trigger unit receives and determines whether a sound input meets certain criteria (e.g., listening mode). In a second mode, the voice trigger unit 902 does not receive and / or process sound input (e.g., standby mode). The electronic device 900 also includes a processing unit 906 coupled to the voice trigger unit 902. In some implementations, the processing unit 906 includes an environment detector 908 that may include and / or interface with one or more sensors (e.g., including a microphone, a camera, an accelerometer, a gyroscope, etc.) and a mode switching unit 910. In some implementations, the processing unit 906 is configured to determine whether the electronic device is within a substantially enclosed space (e.g., via the environment detection unit 908) by detecting that one or more of the microphone and camera of the electronic device are blocked, and to switch the voice trigger to a second mode (e.g., via the mode switching unit 910) upon determining that the electronic device is within the substantially enclosed space.

[0141] In some implementations, the processing unit is configured to determine (e.g., via the environment detection unit 908) whether the electronic device is in a predetermined orientation, and upon determining that the electronic device is in the predetermined orientation, enable (e.g., via the mode switching unit 910) a predetermined mode of the voice trigger.

[0142] According to some implementations, FIG. 10 shows a functional block diagram of an electronic device 1000 configured according to the principles of the present invention as described above. The functional blocks of the device can be implemented by hardware, software, or a combination of hardware and software to execute the principles of the present invention. It is understood by those skilled in the art that the functional blocks described in FIG. 10 can be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description in this specification may support any possible combination or division, or further definition of the functional blocks described in this specification.

[0143] As shown in FIG. 10, the electronic device 1000 includes a voice trigger unit 1002. The voice trigger unit 1002 can be operated in a variety of different modes. In a first mode, the voice trigger unit receives and determines whether a sound input meets certain criteria (e.g., listening mode). In a second mode, the voice trigger unit 1002 does not receive and / or process the sound input (e.g., standby mode). The electronic device 1000 also includes a processing unit 1006 coupled to the voice trigger unit 1002. In some implementations, the processing unit 1006 includes an environment detector 1008, which may include and / or interface with a microphone and / or camera, and a mode switching unit 1010.

[0144] The processing unit 1006 is configured to determine whether the electronic device is in a substantially enclosed space (e.g., via the environment detection unit 1008) by detecting that one or more of the microphone and camera of the electronic device are blocked, and to switch the voice trigger to a second mode (e.g., via the mode switching unit 1010) upon determining that the electronic device is in a substantially enclosed space. The above description has been described with reference to specific embodiments for illustrative purposes. However, the exemplary description above is not intended to be exhaustive or to limit the disclosed embodiments to the precise forms. Many modifications and variations are possible in light of the above teachings. The embodiments have been selected and described in order to best explain the principles and practical applications of the disclosed ideas, thereby enabling those skilled in the art to best utilize the same with various modifications suited to the particular applications envisaged.

[0145] It should be understood that terms such as "first", "second", etc. may be used herein to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first sound detector can be referred to as a second sound detector, and similarly, a second sound detector can be referred to as a first sound detector, without changing the meaning of the description, so long as the name is consistently changed for all occurrences of "first sound detector" and consistently changed for all occurrences of "second sound detector". The first sound detector and the second sound detector are both sound detectors, but are not the same sound detector.

[0146] The terms used herein are for the purpose of describing particular embodiments and are not intended to limit the scope of the claims. As used in the description of the illustrated embodiments and in the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, it is also to be understood that the term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items. It is further to be understood that the terms "comprises" and / or "comprising", as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "if" can be interpreted to mean "when" or "upon" or "upon determining" or "in accordance with determining" or "upon detecting" that the previously stated condition is true, depending on the context. Similarly, the phrases "if it is determined that "[the foregoing condition is true]" or "when "[the foregoing condition is true]" or "when "[the foregoing condition is true]" may be interpreted to mean "upon determining," "upon determining of," "in response to determining," "according to determining," "upon detecting," or "in response to detecting" that the foregoing condition is true.

Claims

1. A method for operating a voice trigger, the method being performed on an electronic device including one or more processors and a memory storing instructions for execution by the one or more processors, the method comprising: operating the voice trigger in a first mode; determining, while operating the voice trigger in the first mode, that the electronic device is within physical proximity of a second electronic device, the second electronic device operating a second voice trigger; in response to determining that the electronic device is within physical proximity of the second electronic device, switching the voice trigger to a second mode.

2. The method of claim 1, wherein the electronic device includes a first microphone, the method comprising: monitoring audio input using the first microphone while operating the voice trigger in the first mode and while the electronic device is coupled to an external electronic device that includes a second microphone; in response to determining that the audio input includes predetermined content, Launching voice-based services; and causing the second microphone to monitor a second audio input; ceasing to monitor the audio input using the first microphone.

3. The method of claim 1, wherein the electronic device includes a first microphone, the method comprising: While operating the voice trigger in the first mode, monitoring the first audio input using the first microphone to determine whether the first audio input includes predetermined content; while the electronic device is coupled to an external electronic device that includes a second microphone; Following a determination that a specified condition has been met, causing the second microphone to monitor the second audio input to determine whether the second audio input includes the predetermined content; ceasing to monitor the first audio input using the first microphone.

4. A method as described in claim 3, wherein determining that the predetermined condition is satisfied includes determining that the external electronic device is currently worn by a user.

5. The method of claim 4, wherein determining that the external electronic device is currently worn by the user comprises: determining a difference between a movement of the electronic device and a movement of the external electronic device.

6. The method according to claim 3, upon determining that the second audio input includes the predetermined content, The method further comprising initiating a voice-based service operating on the electronic device.

7. The method of claim 6, wherein the voice-based service includes a digital assistant.

8. The method of claim 3, wherein the external electronic device includes a headset.

9. The method of claim 3, wherein the predetermined content includes one or more words.

10. The method of claim 1, wherein the first mode includes a listening mode.

11. The method of claim 1, wherein the second mode includes a standby mode.

12. The method of claim 1, wherein determining that the electronic device is within physical proximity of the second electronic device comprises: determining physical proximity to the second electronic device using RFID communications.

13. The method of claim 1, wherein determining that the electronic device is within physical proximity of the second electronic device comprises: determining physical proximity to the second electronic device using proximity communication.

14. A computer program causing a computer to carry out a method according to any one of claims 1 to 13.

15. An electronic device comprising: A memory for storing a computer program according to claim 14; and one or more processors capable of executing the computer programs stored in the memory.

Citation Information

Patent Citations

  • Voice input device, voice recognition system and voice recognition method

    JP2010217754A

  • JPP4319573B

  • Mobile personal audio device

    US20090318198A1