Voice trigger for digital assistant
A low-power voice trigger system for digital assistants uses sound detectors to recognize specific phrases, addressing the need for hands-free operation and reducing power consumption by employing a duty cycle and multiple detectors.
Patent Information
- Application Number
- JP2025088156
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2013-02-07
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-02
AI Technical Summary
Existing voice-based digital assistants require tactile input to initiate, consuming power and inhibiting hands-free operation, and continuous monitoring for voice triggers is power-intensive.
Implement a low-power voice trigger system using sound detectors to recognize specific words or phrases without continuous processing, employing a duty cycle and multiple detectors to reduce power consumption.
Enables hands-free operation with reduced power consumption by using low-power sound detectors to activate voice-based assistants with specific triggers, optimizing battery life in portable devices.
Smart Images

Figure 2025128187000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 61 / 762,260, entitled "VOICE TRIGGER FOR A DIGITAL ASSISTANT," filed February 7, 2013, and is hereby incorporated by reference in its entirety for all purposes.
[0002] <Technical field> The disclosed embodiments relate generally to digital assistants, and more particularly to methods and systems for voice triggers for digital assistants. [Background technology]
[0003] In recent years, voice-based digital assistants, such as Apple's SIRI®, have been introduced to the market to handle various tasks, such as web searching and navigation. One advantage of such voice-based digital assistants is that a user can interact with the device in a hands-free manner without having to manipulate or look at the device. Hands-free operation can be particularly useful when a person cannot or should not physically operate a device, such as while driving. However, to initiate a voice-based assistant, a user typically must press a button or select an icon on a touchscreen. This tactile input inhibits the hands-free experience. Accordingly, it would be advantageous to provide a method and system for enabling a voice-based digital assistant (or other speech-based service) using voice input or signals rather than tactile input.
[0004] Activating a voice-based assistant using voice input requires monitoring the audio channel to detect voice input. This monitoring consumes power, which is a limited resource on the battery-dependent handheld or portable devices on which such voice-based digital assistants often run. Therefore, it would be beneficial to provide an energy-efficient voice trigger that can be used to initiate voice-based and / or speech-based services on the device. Summary of the Invention
[0005] Accordingly, there is a need for a low-power voice trigger that can provide "always listening" voice trigger functionality without excessively consuming limited power resources. The embodiments described below provide systems and methods for using voice triggers in electronic devices to initiate a voice-based assistant. Interaction with a voice-based digital assistant (or other speech-based service, such as a speech-to-text transcription service) often begins when a user presses an affordance (e.g., a button or icon) on the device to activate the digital assistant, and the device then provides some indication to the user that the digital assistant is active and listening, such as a light, a sound (e.g., a beep), or a spoken output (e.g., "What can I do for you?"). As described herein, voice triggers can also be implemented to activate in response to specific, predetermined words, phrases, or sounds without requiring physical interaction by the user. For example, a user may be able to activate the SIRI digital assistant on an iPhone® (both offered by Apple Inc., the assignee of the present application) by speaking the phrase "call SIRI." In response, the device emits a beep, tone, or speech output (e.g., "What can I do for you?") to indicate to the user that listening mode is active. In response, the user can begin interacting with the digital assistant without having to physically touch the device that provides the digital assistant functionality.
[0006] One technique for initiating a speech-based service with a voice trigger is to have the speech-based service continuously listen for a predetermined trigger word, phrase, or sound (any of which may be referred to herein as a “trigger sound”). However, continuously operating a speech-based service (e.g., a voice-based digital assistant) requires significant voice processing and battery power. Various techniques may be employed to reduce power consumption in providing voice trigger functionality. In some implementations, the main processor (i.e., the “application processor”) of the electronic device is maintained in a low-power or no-power state while one or more low-power sound detectors are maintained active (e.g., because they are independent of the application processor). (When in a low-power or no-power state, the application processor or any other processor, program, or module may be described as being disabled or in standby mode.) For example, even when the application processor is disabled, a low-power sound detector is used to monitor the audio channel for a trigger sound. This sound detector is sometimes referred to herein as a trigger sound detector. In some implementations, the low-power sound detector is configured to detect specific sounds, phonemes, and / or words. Trigger sound detectors (including hardware and / or software components) are designed to recognize distinctive words, sounds, or phrases, but are not typically able to provide or are not optimized for full speech-to-text functionality because such a task requires significant computational and power resources. Thus, in some implementations, the trigger sound detector recognizes whether the voice input contains a predetermined pattern (e.g., a sonic pattern matching the words "to SIRI"), but is not able to (or configured to) convert the voice input to text or recognize many other words. When a trigger sound is detected, the digital assistant is subsequently brought out of standby mode so that the user can provide a voice command.
[0007] In some implementations, the trigger sound detector is configured to detect a variety of different trigger sounds, such as a set of words, phrases, sounds, and / or combinations thereof. The user can then use any of these sounds to initiate a speech-based service. In one example, the voice trigger is pre-configured to respond to the phrases "Call SIRI," "Wake SIRI," "Call digital assistant," or "Hello, HAL, can you hear me, HAL?" In some implementations, the user is required to select one of the pre-configured trigger sounds as the single trigger sound. In some implementations, the user selects a subset of the pre-configured trigger sounds so that the user can initiate a speech-based service with different trigger sounds. In some implementations, all of the pre-configured trigger sounds remain valid trigger sounds.
[0008] In some implementations, a separate sound detector is used, and the trigger sound detector may remain in a low-power or no-power mode much of the time. For example, a different type of sound detector (e.g., one that uses less power than the trigger sound detector) is used to monitor the audio channel and determine whether the sound input corresponds to a particular type of sound. Sounds are classified into different "types" based on certain distinguishable sound characteristics. For example, sounds belonging to the "human voice" type have particular spectral content, periodicity, fundamental frequency, etc. Other types of sounds (e.g., whistles, applause, etc.) have different characteristics. The different types of sounds are identified using voice processing and / or signal processing techniques, as described herein. This sound detector is sometimes referred to herein as a "sound type detector." For example, if the predetermined trigger phrase is "To SIRI," the sound type detector determines whether the input approximately corresponds to a human speaking voice. If the trigger sound is a non-voiced sound, such as a whistle, the sound type detector determines whether the sound input approximately corresponds to a whistle. When the appropriate type of sound is detected, the sound type detector activates the trigger sound detector to further process and / or analyze the sound. Because the sound type detector requires less power than the trigger sound detector (e.g., because it uses less power-demanding circuitry and / or more efficient audio processing algorithms than the trigger sound detector), the voice trigger function consumes less power than the trigger sound detector alone.
[0009] In some implementations, an additional sound detector is used, and both the sound type detector and the trigger sound detector may be maintained in a low-power or no-power mode much of the time. For example, a sound detector using less power than the sound type detector may be used to monitor an audio channel to determine whether the sound input meets a predetermined condition, such as an amplitude threshold (e.g., volume). This sound detector may also be referred to herein as a noise detector. When the noise detector detects a sound meeting the predetermined threshold, it activates the sound type detector to further process and / or analyze the sound. Because the noise detector requires less power than either the sound type detector or the trigger sound detector (e.g., because it uses less power-demanding circuitry and / or more efficient audio processing algorithms), the voice trigger function consumes less power than the combination of the sound type detector and the trigger sound detector without the noise detector.
[0010] In some embodiments, any one or more of the sound detectors described above are operated according to a duty cycle that cycles between an “on” and “off” state. This further helps reduce power consumption of the voice trigger. For example, in some embodiments, the noise detector is “on” (i.e., actively monitoring the audio channel) for 10 milliseconds, followed by “off” for 90 milliseconds. In this way, the noise detector is “off” 90% of the time while still effectively providing continuous noise detection functionality. In some embodiments, the on and off durations for the sound detectors are selected so that all of the detectors are enabled while trigger sounds are still being input. For example, for a trigger phrase “to SIRI,” the sound detectors may be configured so that no matter where in the duty cycle the trigger phrase begins, the trigger sound detector is enabled in time to analyze a sufficient amount of input. For example, the trigger sound detector is enabled in time to receive, process, and analyze enough of the sound “to IRI” to determine that the sound matches the trigger phrase. In some implementations, the sound input is stored in memory as it is received and passed to an upstream detector so that the majority of the sound input can be analyzed. Accordingly, even if the trigger sound detector is not activated until after the trigger phrase is spoken, the entire recorded trigger phrase can still be analyzed.
[0011] Some embodiments provide a method of operating a voice trigger. The method is implemented in an electronic device including one or more processors and a memory storing instructions executable by the one or more processors. The method includes receiving a sound input. The method further includes determining whether at least a portion of the sound input corresponds to a predetermined type of sound. The method further includes, upon determining that at least a portion of the sound input corresponds to the predetermined type, determining whether the sound input includes predetermined content. The method further includes, upon determining that the sound input includes the predetermined content, initiating a speech-based service. In some embodiments, the speech-based service is a voice-based digital assistant. In some embodiments, the speech-based service is a dictation service.
[0012] In some embodiments, determining whether the sound input corresponds to a predetermined type of sound is performed by a first sound detector, and determining whether the sound input includes predetermined content is performed by a second sound detector. In some embodiments, the first sound detector consumes less power during operation than the second sound detector. In some embodiments, the first sound detector performs a frequency domain analysis of the sound input. In some embodiments, determining whether the sound input corresponds to a predetermined type of sound is performed upon determining that the sound input satisfies a predetermined condition (e.g., as determined by a third sound detector, described below).
[0013] In some implementations, the first sound detector periodically monitors the audio channel according to a duty cycle, which in some implementations includes an on-time of about 20 milliseconds and an off-time of about 100 milliseconds.
[0014] In some embodiments, the predetermined type is a human voice and the predetermined content is one or more words. In some embodiments, determining whether at least a portion of the sound input corresponds to a predetermined type of sound includes determining whether at least a portion of the sound input includes a frequency characteristic of a human voice.
[0015] In some embodiments, the second sound detector is activated in response to the first sound detector determining that the sound input corresponds to a predetermined type. In some embodiments, the second sound detector is operated for at least a predetermined time period after the first sound detector determines that the sound input corresponds to a predetermined type. In some embodiments, the predetermined time period corresponds to a duration of the predetermined content.
[0016] In some embodiments, the predetermined content is one or more predetermined phonemes. In some embodiments, the one or more predetermined phonemes comprise at least one word.
[0017] In some embodiments, the method includes determining whether the sound input satisfies a predetermined condition before determining whether the sound input corresponds to a predetermined type of sound. In some embodiments, the predetermined condition is an amplitude threshold. In some embodiments, determining whether the sound input satisfies the predetermined condition is performed by a third sound detector, the third sound detector consuming less power during operation than the first sound detector. In some embodiments, the third sound detector periodically monitors the audio channel according to a duty cycle. In some embodiments, the duty cycle includes an on-time of about 20 milliseconds and an off-time of about 500 milliseconds. In some embodiments, the third sound detector performs a time-domain analysis of the sound input.
[0018] In some implementations, the method includes storing at least a portion of the sound input in a memory, and providing the portion of the sound input to the speech-based service when the speech-based service is initiated. In some implementations, the portion of the sound input is stored in the memory using direct memory access.
[0019] In some embodiments, the method includes determining whether the sound input corresponds to a voice of the specified user. In some embodiments, the speech-based service is initiated upon determining that the sound input includes predetermined content and that the sound input corresponds to a voice of the specified user. In some embodiments, the speech-based service is initiated in a restricted access mode upon determining that the sound input includes predetermined content and that the sound input does not correspond to a voice of the specified user. In some embodiments, the method includes outputting an audio prompt including the name of the specified user upon determining that the sound input corresponds to a voice of the specified user.
[0020] In some embodiments, determining whether the sound input includes the predetermined content includes comparing a representation of the sound input to a reference representation, and determining that the sound input includes the predetermined content if the representation of the sound input matches the reference representation. In some embodiments, a match is determined if the representation of the sound input matches the reference representation with a predetermined confidence value. In some embodiments, the method includes receiving a plurality of sound inputs including the sound input, and iteratively adjusting the reference representation using each one of the plurality of sound inputs in response to determining that each sound input includes the predetermined content.
[0021] In some embodiments, the method includes determining whether the electronic device is in a predetermined orientation and, upon determining that the electronic device is in the predetermined orientation, enabling a predetermined mode of the voice trigger. In some embodiments, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing downward, and the predetermined mode is a standby mode. In some embodiments, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing upward, and the predetermined mode is a listening mode.
[0022] Some embodiments provide a method of operating a voice trigger. The method is implemented in an electronic device including one or more processors and a memory storing instructions executable by the one or more processors. The method includes operating the voice trigger in a first mode. The method further includes determining whether the electronic device is within a substantially enclosed space by detecting that one or more of a microphone and a camera of the electronic device are occluded. The method further includes switching the voice trigger to a second mode upon determining that the electronic device is within the substantially enclosed space. In some embodiments, the second mode is a standby mode.
[0023] Some embodiments provide a method of operating a voice trigger. The method is implemented in an electronic device including one or more processors and a memory storing instructions executable by the one or more processors. The method includes determining whether the electronic device is in a predetermined orientation and, upon determining that the electronic device is in the predetermined orientation, enabling a predetermined mode of the voice trigger. In some embodiments, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing downward, and the predetermined mode is a standby mode. In some embodiments, the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing upward, and the predetermined mode is a listening mode.
[0024] In some embodiments, the electronic device includes a receiving unit configured to receive an audio input and a processing unit coupled to the receiving unit. The processing unit is configured to determine whether at least a portion of the audio input corresponds to a predetermined type of sound, determine whether the audio input includes predetermined content upon determining that at least a portion of the audio input corresponds to the predetermined type, and initiate a speech-based service upon determining that the audio input includes the predetermined content. In some embodiments, the processing unit is further configured to determine whether the audio input satisfies a predetermined condition before determining whether the audio input corresponds to the predetermined type of sound. In some embodiments, the processing unit is further configured to determine whether the audio input corresponds to the voice of a specific user.
[0025] In some embodiments, the electronic device includes a voice trigger unit configured to operate a voice trigger in a first mode of a plurality of modes and a processing unit coupled to the voice trigger unit. In some embodiments, the processing unit is configured to determine whether the electronic device is in a substantially enclosed space by detecting that one or more of a microphone and a camera of the electronic device are occluded, and to switch the voice trigger to a second mode upon determining that the electronic device is in the substantially enclosed space. In some embodiments, the processing unit is configured to determine whether the electronic device is in a predetermined orientation, and to enable the predetermined mode of the voice trigger upon determining that the electronic device is in the predetermined orientation.
[0026] According to some embodiments, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided that stores one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for performing any of the methods described herein.
[0027] According to some embodiments, there is provided an electronic device (eg, a portable electronic device) that includes means for performing any of the methods described herein.
[0028] According to some embodiments, there is provided an electronic device (eg, a portable electronic device) that includes a processing unit configured to perform any of the methods described herein.
[0029] According to some embodiments, an electronic device (e.g., a portable electronic device) is provided that includes one or more processors and memory that stores one or more programs that are executed by the one or more processors, the one or more programs including instructions for performing any of the methods described herein.
[0030] According to some embodiments, there is provided an information processing device for use in an electronic device, the information processing device including means for performing any of the methods described herein. [Brief explanation of the drawings]
[0031] [Figure 1] FIG. 1 is a block diagram illustrating an environment in which a digital assistant operates, according to some embodiments.
[0032] [Figure 2] FIG. 1 is a block diagram illustrating a digital assistant client system according to some embodiments.
[0033] [Figure 3A] FIG. 1 is a block diagram illustrating a standalone digital assistant system or a digital assistant server system according to some embodiments.
[0034] [Figure 3B] FIG. 3B is a block diagram illustrating the functionality of the digital assistant shown in FIG. 3A, according to some embodiments.
[0035] [Figure 3C] FIG. 1 is a network diagram illustrating a portion of an ontology, according to some embodiments.
[0036] [Figure 4] FIG. 1 is a block diagram illustrating components of a voice trigger system, according to some embodiments.
[0037] [Figure 5] 1 is a flowchart illustrating a method for operating a voice trigger system according to some embodiments. [Figure 6] 1 is a flowchart illustrating a method for operating a voice trigger system according to some embodiments. [Figure 7] 1 is a flowchart illustrating a method for operating a voice trigger system according to some embodiments.
[0038] [Figure 8] FIG. 1 is a functional block diagram of an electronic device in accordance with some embodiments. [Figure 9] FIG. 1 is a functional block diagram of an electronic device in accordance with some embodiments.
[0039] Like reference numbers refer to corresponding parts throughout the drawings. DETAILED DESCRIPTION OF THE INVENTION
[0040] Figure 1 is a block diagram of a digital assistant operating environment 100 according to some embodiments. The terms "digital assistant," "virtual assistant," "intelligent automated assistant," "voice-based digital assistant," or "automated digital assistant" refer to any information processing system that interprets spoken and / or textual natural language input to infer a user's intent (e.g., identify a task type corresponding to the natural language input) and perform an action based on the inferred user intent (e.g., perform a task corresponding to the identified task type). For example, to act based on the inferred user intent, the system can perform one or more of the following: identify a task flow having steps and parameters designed to fulfill the inferred user intent (e.g., identify a task type), input specific requests from the inferred user intent into the task flow, execute the task flow by invoking a program, method, service, API, or the like (e.g., send a request to a service provider), and generate an output response to the user in an audible (e.g., speech) and / or visual form.
[0041] Specifically, once initiated, the digital assistant system can accept user requests, at least in part, in the form of natural language commands, requests, statements, narratives, and / or queries. Generally, user requests seek either an informational answer or task performance by the digital assistant system. A satisfactory response to a user request typically involves providing the requested informational answer, performing the requested task, or some combination of the two. For example, a user may ask the digital assistant system a question such as, "Where am I right now?" Based on the user's current location, the digital assistant may respond, "You're near the West Gate in Central Park." A user can also request task performance by stating, for example, "Please invite my friends to my girlfriend's birthday party next week." In response, the digital assistant may confirm the request by generating a voice output of, "Yes, right away," and then send appropriate calendar invitations from the user's email address to each of the user's friends listed in the user's electronic address book or contact list. There are many other ways to interact with a digital assistant to request information or perform various tasks. In addition to providing verbal responses and taking programmed actions, digital assistants can also provide other visual or audio forms of responses (e.g., as text, alerts, music, videos, animations, etc.).
[0042] As shown in FIG. 1 , in some embodiments, the digital assistant system is implemented according to a client-server model. The digital assistant system includes a client-side portion (e.g., 102a and 102b) (hereinafter, "digital assistant (DA) client 102") that runs on user devices (e.g., 104a and 104b) and a server-side portion 106 (hereinafter, "digital assistant (DA) server 106") that runs on a server system 108. The DA client 102 communicates with the DA server 106 over one or more networks 110. The DA client 102 provides client-side functionality, such as user-responsive input and output processing and communication with the DA server 106. The DA server 106 provides server-side functionality for any number of DA clients 102, each residing on a respective user device 104 (also referred to as a client device or electronic device).
[0043] In some embodiments, the DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, an I / O interface 118 to external services, a photo and tag database 130, and a photo tag module 132. The client-facing I / O interface facilitates client-facing input and output processing for the digital assistant server 106. The one or more processing modules 114 utilize the data and models 116 to determine user intent based on natural language input and perform tasks based on the estimated user intent. The photo and tag database 130 stores digital photo fingerprints and, optionally, the digital photos themselves, as well as tags associated with the digital photos. The photo tag module 132 creates and stores tags associated with photos and / or fingerprints, automatically tags photos, and links tags to locations within photos.
[0044] In some implementations, the DA server 106 communicates with external services 120 (e.g., navigation service(s) 122-1, messaging service(s) 122-2, information service(s) 122-3, calendar service 122-4, phone service 122-5, photo service(s) 122-6, etc.) through network(s) 110 to complete tasks or obtain information. An I / O interface 118 to external services facilitates such communication.
[0045] Examples of user equipment 104 include, but are not limited to, a handheld computer, a wireless personal digital assistant (PDA), a tablet computer, a laptop computer, a desktop computer, a cellular telephone, a smart phone, an enhanced general packet radio service (EGPRS) mobile phone, a media player, a navigation device, a game console, a television, a remote control, or a combination of any two or more of these data processing devices, or any other suitable data processing device. Further details regarding user equipment 104 are provided with respect to the exemplary user equipment 104 shown in FIG. 2.
[0046] Examples of communication network(s) 110 include local area networks (LANs) and wide area networks (WANs), such as the Internet. Communication network(s) 110 may be implemented using any known network protocol, including various wired or wireless protocols, such as Ethernet, Universal Serial Bus (USB), FIREWIRE, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wi-Fi, voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.
[0047] The server system 108 may be implemented on at least one data processing device and / or a distributed network of computers. In some implementations, the server system 108 also utilizes various virtual appliances and / or the services of third-party service providers (e.g., third-party cloud service providers) to provide the underlying computing and / or infrastructure resources of the server system 108.
[0048] The digital assistant system shown in FIG. 1 includes both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), but in some embodiments, the digital assistant system refers only to the server-side portion (e.g., DA server 106). In some embodiments, the digital assistant's functionality can be implemented as a standalone application installed on user equipment. In addition, the distribution of functionality between the client and server portions of the digital assistant can vary depending on the embodiment. For example, in some embodiments, DA client 102 is a thin client that provides only user-facing input and output processing functionality and delegates all other digital assistant functionality to DA server 106. For example, in some embodiments, DA client 102 is configured to perform or assist one or more functions of DA server 106.
[0049] 2 is a block diagram of user equipment 104, according to some embodiments. User equipment 104 includes a memory interface 202, one or more processors 204, and a peripherals interface 206. The various components within user equipment 104 are coupled by one or more communication buses or signal lines. User equipment 104 includes various sensors, subsystems, and peripherals coupled to peripherals interface 206. The sensors, subsystems, and peripherals collect information and / or facilitate various functions of user equipment 104.
[0050] For example, in some implementations, a motion sensor 210 (e.g., an accelerometer), a light sensor 212, a GPS receiver 213, a temperature sensor, and a proximity sensor 214 are coupled to the peripherals interface 206 to facilitate orientation, light, and proximity sensing functions. In some implementations, other sensors 216, such as biometric sensors, barometers, etc., are connected to the peripherals interface 206 to facilitate related functions.
[0051] In some implementations, user equipment 104 includes a camera subsystem 220 coupled to peripherals interface 206. In some implementations, an optical sensor 222 in camera subsystem 220 facilitates camera functions such as taking pictures and recording video clips. In some implementations, user equipment 104 includes one or more wired and / or wireless communication subsystems 224 that provide communication functions. Communication subsystem 224 typically includes various communication ports, radio frequency receivers and transmitters, and / or optical (e.g., infrared) receivers and transmitters. In some implementations, user equipment 104 includes an audio subsystem 226 coupled to one or more speakers 228 and one or more microphones 230 to facilitate voice-enabled functions such as voice recognition, voice response, digital recording, and telephone functions. In some implementations, audio subsystem 226 is coupled to voice trigger system 400. In some implementations, voice trigger system 400 and / or audio subsystem 226 include low-power audio circuitry and / or programs (i.e., including hardware and / or software) for receiving and / or analyzing sound input, including, for example, one or more analog-to-digital converters, digital signal processors (DSPs), sound detectors, memory buffers, codecs, etc. In some implementations, the low-power audio circuitry (alone or in addition to other components of user equipment 104) provides voice (or sound) trigger functionality for one or more aspects of user equipment 104, such as a voice-based digital assistant or other speech-based service. In some implementations, the low-power audio circuitry provides voice trigger functionality even when other components of user equipment 104, such as processor(s) 204, I / O subsystem 240, memory 250, etc., are turned off and / or in standby mode. This voice trigger system 400 is described in further detail with respect to FIG. 4.
[0052] In some implementations, I / O subsystem 240 is also coupled to peripherals interface 206. In some implementations, user device 104 includes a touchscreen 246, and I / O subsystem 240 includes a touchscreen controller 242 coupled to touchscreen 246. If user device 104 includes touchscreen 246 and touchscreen controller 242, touchscreen 246 and touchscreen controller 242 are typically configured to detect contact and movement or disruption using any of a number of touch-sensing technologies, such as, for example, capacitive, resistive, infrared, surface ultrasonic technology, proximity sensor arrays, and the like. In some implementations, user device 104 includes a display that does not include a touch-sensitive surface. In some implementations, user device 104 includes a separate touch-sensitive surface. In some implementations, user device 104 includes other input controller(s) 244. If the user device 104 includes other input controller(s) 244, the other input controller(s) 244 are typically coupled to other input / control devices 248, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointer devices such as styluses.
[0053] Memory interface 202 is coupled to memory 250. In some implementations, memory 250 includes a persistent computer-readable medium, such as high-speed random-access memory and / or non-volatile memory (e.g., one or more magnetic disk storage devices, one or more flash memory devices, one or more optical storage devices, and / or other non-volatile solid-state storage devices). In some implementations, memory 250 stores an operating system 252, a communications module 254, a graphical user interface module 256, a sensor processing module 258, a telephony module 260, and applications 262, or a subset or superset thereof. Operating system 252 includes instructions for handling basic system services and for performing hardware-dependent tasks. Communications module 254 facilitates communication with one or more additional devices, one or more computers, and / or one or more servers. Graphical user interface module 256 facilitates graphic user interface processing. Sensor processing module 258 facilitates sensor-related processing and functions (e.g., processing audio input received using one or more microphones 228). Telephony module 260 facilitates telephony-related processes and functions. Application module 262 facilitates various functions of a user application, such as electronic messaging, web browsing, media processing, navigation, imaging, and / or other processes and functions. In some implementations, user equipment 104 stores in memory 250 one or more software applications 270-1 and 270-2, each associated with at least one of the external service providers.
[0054] As mentioned above, in some implementations, memory 250 also stores client-side digital assistant instructions (e.g., in digital assistant client module 264) and various user data 266 (e.g., user-specific vocabulary data, preference data, and / or other data such as the user's electronic address book or contact list, to-do list, shopping list, etc.) to provide the client-side functionality of the digital assistant.
[0055] In some implementations, digital assistant client module 264 can accept voice input, text input, touch input, and / or gesture input through various user interfaces (e.g., I / O subsystem 244) of user equipment 104. Digital assistant client module 264 can also provide output in audio, visual, and / or tactile forms. For example, output can be provided as voice, sound, alerts, text messages, menus, graphics, video, animations, vibrations, and / or combinations of two or more of the above. In operation, digital assistant client module 264 communicates with a digital assistant server (e.g., digital assistant server 106, FIG. 1) using communication subsystem 224.
[0056] In some embodiments, digital assistant client module 264 uses various sensors, subsystems, and peripherals to collect additional information from the surrounding environment of user device 104 to establish a context associated with the user input. In some embodiments, digital assistant client module 264 provides the context information, or a subset thereof, along with the user input to the digital assistant server (e.g., digital assistant server 106, FIG. 1) to help infer the user's intent.
[0057] In some embodiments, context information that may accompany a user input includes sensor information, such as ambient lighting, ambient noise, ambient temperature, images or video, etc. In some embodiments, context information also includes the physical state of the device, such as device orientation, device location, device temperature, power level, speed, acceleration, movement patterns, cellular signal strength, etc. In some embodiments, information related to the software state of user device 106, such as user device 104's running processes, installed programs, past and present network activity, background services, error logs, resource usage, etc., is also provided to the digital assistant server (e.g., digital assistant server 106, FIG. 1) as context information associated with the user input.
[0058] In some embodiments, the DA client module 264 selectively provides information stored on the user device 104 (e.g., at least a portion of the user data 266) in response to a request from the digital assistant server. In some embodiments, the digital assistant client module 264 also elicits additional input from the user via a natural language dialog or other user interface in response to a request by the digital assistant server 106 (FIG. 1). The digital assistant client module 264 passes the additional input to the digital assistant server 106 to assist the digital assistant server 106 in estimating and / or achieving the user intent expressed in the user request.
[0059] In some implementations, memory 250 may include additional or fewer instructions. Furthermore, various functions of user equipment 104 may be implemented in hardware and / or firmware, including in the form of one or more signal processing and / or application specific integrated circuits, and therefore user equipment 104 need not include all of the modules and applications shown in FIG.
[0060] FIG. 3A is a block diagram of an exemplary digital assistant system 300 (also referred to as a digital assistant) according to some embodiments. In some embodiments, digital assistant system 300 is implemented on a standalone computer system. In some embodiments, digital assistant system 300 is distributed across multiple computers. In some embodiments, some of the digital assistant's modules and functionality are divided into a server portion and a client portion. The client portion resides on user equipment (e.g., user equipment 104) and communicates with the server portion (e.g., server system 108) over one or more networks, for example, as shown in FIG. 1. In some embodiments, digital assistant system 300 is an embodiment of server system 108 (and / or digital assistant server 106) shown in FIG. 1. In some embodiments, digital assistant system 300 is implemented within user equipment (e.g., user equipment 104, FIG. 1), thereby eliminating the need for a client-server system. It should be noted that digital assistant system 300 is merely one example of a digital assistant system, and that digital assistant system 300 may have more or fewer components than shown, may combine two or more components, or may have a different configuration or arrangement of components. The various components shown in FIG. 3A may be implemented in hardware, software, firmware, or a combination thereof, including one or more signal processing and / or application specific integrated circuits.
[0061] Digital assistant system 300 includes memory 302, one or more processors 304, an input / output (I / O) interface 306, and a network communication interface 308. These components communicate with each other through one or more communication buses or signal lines 310.
[0062] In some implementations, memory 302 includes a persistent computer-readable medium, such as high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., one or more magnetic disk storage devices, one or more flash memory devices, one or more optical storage devices, and / or other non-volatile solid-state memory devices).
[0063] I / O interface 306 connects input / output devices 316 of digital assistant system 300, such as a display, keyboard, touchscreen, and microphone, to user interface module 322. I / O interface 306 cooperates with user interface module 322 to receive user inputs (e.g., voice input, keyboard input, touch input, etc.) and process them appropriately. In some embodiments, when the digital assistant is implemented on a standalone user device, digital assistant system 300 includes any of the components described with respect to user device 104 in FIG. 2 as well as I / O and communication interfaces (e.g., one or more microphones 230). In some embodiments, digital assistant system 300 represents the server portion of a digital assistant implementation and interacts with the user through a client-side portion residing on the user device (e.g., user device 104 shown in FIG. 2).
[0064] In some implementations, network communication interface 308 includes wired communication port(s) 312 and / or wireless transceiver circuitry 314. The wired communication port(s) receive and transmit communication signals via one or more wired interfaces, such as Ethernet, Universal Serial Bus (USB), FIREWIRE®, etc. The wireless circuitry 314 typically receives and transmits RF and / or optical signals to and from communication networks and other communication devices. Wireless communication may use any of a number of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth®, Wi-Fi®, VoIP, Wi-MAX®, or any other suitable communication protocol. The network communication interface 308 enables communication between the digital assistant system 300 and networks such as the Internet, an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN), and other devices.
[0065] In some embodiments, the persistent computer-readable storage medium of memory 302 stores programs, modules, instructions, and data structures, including all or a subset of: an operating system 318, a communications module 320, a user interface module 322, one or more applications 324, and a digital assistant module 326. One or more processors 304 execute these programs, modules, instructions, and read / write from / to the data structures.
[0066] Operating system 318 (e.g., Darwin®, RTXC®, LINUX®, UNIX®, OS X®, iOS®, Windows®, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage control, power management, etc.) and facilitating communication between various hardware, firmware, and software components.
[0067] The communications module 320 facilitates communication between the digital assistant system 300 and other devices through the network communications interface 308. For example, the communications module 320 can communicate with the communications module 254 of the device 104 shown in FIG. 2. The communications module 320 also includes various software components for processing data received by the wireless circuitry 314 and / or the wired communications port 312.
[0068] In some implementations, the user interface module 322 receives commands and / or input from a user via the I / O interface 306 (e.g., from a keyboard, touchscreen, and / or microphone) and provides user interface objects on a display.
[0069] Applications 324 include programs and / or modules configured to be executed by one or more processors 304. For example, if the digital assistant system is implemented on a standalone user device, applications 324 may include user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 300 is implemented on a server farm, applications 324 may include, for example, resource management applications, diagnostic applications, or scheduling applications.
[0070] Memory 302 also stores digital assistant module 326 (i.e., the server portion of the digital assistant). In some implementations, digital assistant module 326 includes the following submodules, or a subset or superset thereof: input / output processing module 328, speech-to-text (STT) processing module 330, natural language processing module 332, dialog flow processing module 334, task flow processing module 336, service processing module 338, and photo module 132. Each of these processing modules has access to one or more of the following data and models of digital assistant 326, or a subset or superset thereof: ontology 360, vocabulary index 344, user data 348, classification module 349, disambiguation module 350, task flow model 354, service model 356, photo tagging module 358, search module 360, and local tag / photo storage 362.
[0071] In some embodiments, using processing modules (e.g., input / output processing module 328, STT processing module 330, natural language processing module 332, dialog flow processing module 334, task flow processing module 336, and / or service processing module 338), data, and models implemented within digital assistant module 326, digital assistant system 300 does at least some of the following: identify the user's intent expressed in natural language input received from the user, actively elicit and obtain the information necessary to fully infer the user's intent (e.g., by disambiguating words, names, intent, etc.), determine a task flow to satisfy the inferred intent, and execute the task flow to satisfy the inferred intent. In some embodiments, the digital assistant also takes appropriate action when a satisfactory response is not or cannot be provided to the user for various reasons.
[0072] In some embodiments, as described below, digital assistant system 300 processes the natural language input to identify a user's intent and tag digital photos with appropriate information. In some embodiments, digital assistant system 300 also performs other photo-related tasks, such as searching for digital photos using natural language input, automatically tagging photos, etc. As shown in FIG. 3B , in some embodiments, I / O processing module 328 interacts with a user through I / O device 316 of FIG. 3A or with a user device (e.g., user device 104 of FIG. 1) through network communication interface 308 of FIG. 3A to obtain user input (e.g., speech input) and provide a response to the user input. I / O processing module 328 optionally acquires contextual information associated with the user input from the user device upon or immediately after receiving the user input. The contextual information includes user-specific data, vocabulary, and / or preferences associated with the user input. In some embodiments, the context information also includes information about the software and hardware state of the device (e.g., user device 104 in FIG. 1 ) at the time the user request is received and / or information about the user's surrounding environment at the time the user request is received. In some embodiments, I / O processing module 328 also sends follow-up questions to the user regarding the user request and receives responses from the user. In some embodiments, when a user request is received by I / O processing module 328 and the user request includes voice input, I / O processing module 328 forwards the voice input to speech-to-text (STT) processing module 330 for speech-to-text conversion.
[0073] In some implementations, the speech-to-text processing module 330 receives speech input (e.g., a user's utterances captured in an audio recording) through the I / O processing module 328. In some implementations, the speech-to-text processing module 330 uses various acoustic and language models to recognize the speech input as a sequence of phonemes and ultimately as a sequence of words or tokens written in one or more languages. The speech-to-text processing module 330 is implemented using any suitable speech recognition technique, acoustic model, and language model, such as hidden Markov models, dynamic time warping (DTW)-based speech recognition, and other statistical and / or analytical techniques. In some implementations, the speech-to-text processing may be performed at least in part by a third-party service or on the user's device. Once the speech-to-text processing module 330 obtains the results of the speech-to-text processing (e.g., a sequence of words or tokens), it passes the results to the natural language processing module 332 for intent inference. The digital assistant 326's natural language processing module 332 ("natural language processor") takes the string of words or tokens ("token string") generated by the speech-to-text processing module 330 and attempts to associate the token string with one or more "actionable intents" recognized by the digital assistant. As used herein, an "actionable intent" refers to a task that can be performed by the digital assistant 326 and / or the digital assistant system 300 (FIG. 3A) and has an associated task flow implemented in the task flow model 354. The associated task flow is a sequence of programmed actions and steps that the digital assistant system 300 takes to perform the task. The scope of the digital assistant system's capabilities depends on the number and type of task flows implemented and stored in the task flow model 354, or, in other words, the number and type of "actionable intents" recognized by the digital assistant system 300. However, the effectiveness of the digital assistant system 300 also depends on the digital assistant system's ability to infer accurate "actionable intent(s)" from user requests expressed in natural language.
[0074] In some implementations, in addition to the string of words or tokens obtained from speech-to-text processing module 330, natural language processor 332 also receives contextual information associated with the user request (e.g., from I / O processing module 328). Natural language processor 332 optionally uses the contextual information to clarify, complement, and / or further clarify information contained in the string of tokens received from speech-to-text processing module 330. Contextual information includes, for example, user preferences, the state of the hardware and / or software of the user's equipment, sensor information collected before, during, or immediately after the user request, previous interactions (e.g., dialogues) between the digital assistant and the user, and the like.
[0075] In some embodiments, natural language processing is based on ontology 360. Ontology 360 is a hierarchical structure containing multiple nodes, each of which represents either an "actionable intention" or an "attribute" related to one or more of "actionable intentions" or other "attributes." As described above, an "actionable intention" represents a task that digital assistant system 300 is capable of performing (e.g., a task that is "actionable" or can be targeted for performance). An "attribute" represents a parameter associated with a sub-aspect of an actionable intention or another attribute. Links between actionable intention nodes and attribute nodes in ontology 360 define how the parameters represented by the attribute nodes relate to the task represented by the actionable intention node. In some embodiments, ontology 360 is composed of actionable intention nodes and attribute nodes. Within ontology 360, each actionable intention node is linked to one or more attribute nodes directly or via one or more intermediate attribute nodes. Similarly, each attribute node is linked to one or more actionable intention nodes directly or via one or more intermediate attribute nodes. For example, the ontology 360 shown in FIG. 3C includes an actionable intent node, a “Restaurant Reservation” node. The attribute nodes “Restaurant,” “Date / Time” (for reservation), and “Number of Parties” are each directly connected to the “Restaurant Reservation” node (i.e., the actionable intent node). Furthermore, the attribute nodes “Cuisine,” “Price Range,” “Phone Number,” and “Location” are subnodes of the attribute node “Restaurant,” and are each connected to the “Restaurant Reservation” node via the intermediate attribute node “Restaurant.” For another example, the ontology 360 shown in FIG. 3C also includes another actionable intent node, a “Set Reminder” node. The attribute nodes “Date / Time” (for setting a reminder) and “Theme” (for reminders) are each connected to the “Set Reminder” node. Because the attribute node “Date / Time” is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node “Date / Time” is connected to both the “Restaurant Reservation” node and the “Set Reminder” node in the ontology 360.
[0076] An actionable intent node, along with its connected concept nodes, can be described as a "domain." In this description, each domain is associated with a respective actionable intent and refers to the set of nodes (and the relationships between them) associated with a particular actionable intent. For example, ontology 360 shown in FIG. 3C includes an example restaurant reservation domain 362 and an example reminder domain 364 within ontology 360. The restaurant reservation domain includes the actionable intent node "reservation," the attribute nodes "restaurant," "date / time," and "number of parties," and the subattribute nodes "cuisine," "price range," "phone number," and "location." The reminder domain 364 includes the actionable intent node "reminder setting," and the attribute nodes "theme" and "date / time." In some embodiments, ontology 360 is composed of multiple domains. Each domain can share one or more attribute nodes with one or more other domains. For example, the attribute node for "date / time" can be associated with many other domains (e.g., a scheduling domain, a travel reservation domain, a movie ticket domain, etc.) in addition to the restaurant reservation domain 362 and the reminder domain 364. While FIG. 3C shows two example domains within ontology 360, ontology 360 may include other domains (i.e., actionable intents) such as "initiate a call," "get directions," "schedule a meeting," "send a message," and "provide an answer to a question," "tag a photo," etc. For example, the domain for "send a message" is associated with the actionable intent node for "send a message" and can further include attribute nodes such as "recipient(s)," "message type," and "message body." The attribute node "recipient" may be further defined by sub-attribute nodes such as "recipient name" and "message address," for example.
[0077] In some implementations, ontology 360 includes all domains (and therefore actionable intents) that the digital assistant can understand and act upon. In some implementations, ontology 360 may be modified, such as by adding or removing domains or nodes, or by changing the relationships between nodes in ontology 360.
[0078] In some implementations, nodes associated with multiple related actionable intents may be clustered under a "super domain" in ontology 360. For example, a "travel" super domain may include a cluster of travel-related attribute nodes and actionable intent nodes. Actionable intent nodes related to travel may include "book a flight," "book a hotel," "rent a car," "get directions," "find attractions," etc. Actionable intent nodes under the same super domain (e.g., the "travel" super domain) may share many attribute nodes. For example, the actionable intent nodes for "book a flight," "book a hotel," "rent a car," "get directions," and "find attractions" may share one or more of the attribute nodes "departure location," "destination," "departure date / time," "arrival date / time," and "number of participants."
[0079] In some implementations, each node in ontology 360 is associated with a set of words and / or phrases related to the attribute or actionable intent represented by that node. The respective set of words and / or phrases associated with each node is the so-called "vocabulary" associated with that node. The respective set of words and / or phrases associated with each node can be stored in vocabulary index 344 ( FIG. 3B ) in association with the attribute or actionable intent represented by that node. For example, returning to FIG. 3B , the vocabulary associated with a node for the attribute "restaurant" may include words such as "food," "drink," "dish," "hungry," "eat," "pizza," "fast food," and "meal." As another example, the vocabulary associated with a node for the actionable intent "initiate a phone call" may include words and phrases such as "call," "phone," "dial," "ring," "call this number," and "make a call to." The vocabulary index 344 optionally includes words and phrases in different languages. In some embodiments, the natural language processor 332 shown in FIG. 3B receives a token sequence (e.g., a text string) from the speech-to-text processing module 330 and determines which nodes are implied by the words in the token sequence. In some embodiments, if a word or phrase in the token sequence is found to be associated with one or more nodes in the ontology 360 (via the vocabulary index 344), the word or phrase will "trigger" or "activate" those nodes. If multiple nodes are "triggered," the natural language processor 332 will select one of the possible intents as the task (or type of task) that the user intends the digital assistant to perform based on the amount and / or relative importance of the activated nodes. In some embodiments, the domain with the most "triggered" nodes is selected.In some embodiments, the domain with the highest confidence value (e.g., based on the relative importance of its various triggered nodes) is selected. In some embodiments, the domain is selected based on a combination of the number and importance of triggered nodes. In some embodiments, additional factors are also considered when selecting a node, such as whether the digital assistant system 300 has previously correctly interpreted a similar request from the user.
[0080] In some embodiments, digital assistant system 300 also stores names of specific entities in vocabulary index 344. Therefore, when one of these names is detected in a user request, natural language processor 332 can recognize that the name refers to a specific instance of an attribute or subattribute in the ontology. In some embodiments, the names of specific entities are names of businesses, restaurants, people, movies, and the like. In some embodiments, digital assistant system 300 can search for and identify specific entity names from other data sources, such as the user's address book, contact list, movie database, musician database, and / or restaurant database. In some embodiments, when natural language processor 332 identifies a word in a token string as the name of a specific entity (such as a name in the user's address book or contact list), the word is given additional weight in selecting a possible intent in the ontology for the user request. For example, if the word "Mr. Santo" is recognized from a user request, and the last name "Santo" is found in lexical index 344 as one of the contacts in the user's contact list, then the user request likely corresponds to the "send a message" or "initiate a call" domain. As another example, if the word "ABC Cafe" is found in a user request, and the term "ABC Cafe" is found in lexical index 344 as the name of a particular restaurant in the user's city, then the user request likely corresponds to the "restaurant reservation" domain.
[0081] User data 348 includes user-specific information, such as user-specific vocabulary, user preferences, user addresses, the user's default and secondary languages, the user's contact list, and other short-term or long-term information about each user. The natural language processor 332 can use the user-specific information to supplement information contained in the user input to further clarify the user's intent. For example, in response to a user request "invite my friends to my birthday party," the natural language processor 332 can access user data 348 to determine who the "friends" are and when and where the "birthday party" will be held, instead of requiring the user to explicitly provide such information in the user's request.
[0082] In some implementations, the natural language processor 332 includes a classification module 349. In some implementations, the classification module 349 determines whether each of one or more terms in a text string (e.g., corresponding to a voice input associated with a digital photograph) is one of an entity, an action, or a location, as described in more detail below. In some implementations, the classification module 349 classifies each term of the one or more terms as being one of an entity, an action, or a location. Once the natural language processor 332 identifies an actionable intent (or domain) based on the user request, the natural language processor 332 generates a structured query to represent the identified actionable intent. In some implementations, the structured query includes parameters for one or more nodes in the domain of the actionable intent, with at least some of the parameters populated with specific information and requirements specified in the user request. For example, a user may say, "Please make a dinner reservation for me at a sushi restaurant at 7 o'clock." In this case, the natural language processor 332 may be able to accurately identify the actionable intent as "restaurant reservation" based on the user input. According to the ontology, a structured query for the "restaurant reservation" domain may include parameters such as {cuisine}, {time}, {date}, {number of parties}, and the like. Based on information contained in the user's utterance, the natural language processor 332 may generate a partial structured query for the restaurant reservation domain. Here, the partial structured query includes the parameters {cuisine="sushi"} and {time="7:00 PM"}. However, in this example, the user's utterance does not include enough information to complete a structured query associated with the domain. Therefore, other required parameters, such as {number of parties} and {date}, are not specified in the structured query based on the currently available information. In some implementations, the natural language processor 332 adds the received context information to some parameters of the structured query.For example, if a user requests sushi restaurants "near me," the natural language processor 332 may add GPS coordinates from the user device 104 to the {location} parameter in the structured query.
[0083] In some embodiments, the natural language processor 332 passes the structured query (including any completed parameters) to a task flow processing module 336 (“task flow processor”). The task flow processor 336 is configured to receive the structured query from the natural language processor 332, complete the structured query, and perform actions required to “complete” the user's final request. In some embodiments, the various procedures required to complete these tasks are provided in a task flow model 354. In some embodiments, the task flow model 354 includes procedures for obtaining additional information from the user and task flows for performing actions associated with the actionable intent. As described above, to complete the structured query, the task flow processor 336 may need to initiate additional dialogue with the user to obtain additional information and / or disambiguate potentially ambiguous utterances. If such dialogue is required, the task flow processor 336 invokes the dialog processing module 334 (dialog processor) to engage in a dialogue with the user. In some implementations, the dialog processing module 334 determines how (and / or when) to prompt the user for additional information and receives and processes user responses. In some implementations, the dialog processing module 334 provides questions to the user and receives answers from the user through the I / O processing module 328. For example, the dialog processing module 334 presents dialog output to the user via audio and / or visual output and receives input from the user via verbal or physical (e.g., touch gesture) responses. Continuing with the example above, when the task flow processor 336 invokes the dialog processor 334 to determine “number of parties” and “date” information for a structured query associated with the domain “restaurant reservation,” the dialog processor 334 generates questions such as “for how many people?” and “on what date?” to pass to the user.Upon receiving an answer from the user, the dialog processing module 334 passes the information to the task flow processor 336 to add or complete the missing information in the structured query.
[0084] In some cases, task flow processor 336 may receive a structured query with one or more ambiguous attributes. For example, a structured query for the "send a message" domain may indicate that the intended recipient is "Bob," and the user may have multiple contacts named "Bob." Task flow processor 336 will request that dialog processor 334 disambiguate this attribute of the structured query. As a result, dialog processor 334 may ask the user, "Which Bob?" and display (or read out) a list of contacts named "Bob" from which the user can choose.
[0085] In some implementations, dialog processor 334 includes a disambiguation module 350. In some implementations, disambiguation module 350 disambiguates one or more ambiguous terms (e.g., one or more ambiguous terms in a text string corresponding to a voice input associated with a digital photograph). In some implementations, disambiguation module 350 determines that a first term of the one or more terms has multiple possible meanings, prompts a user for additional information about the first term, receives the additional information from the user in response to the prompt, and, in response to the additional information, identifies an entity, action, or location associated with the first term.
[0086] In some implementations, the disambiguation module 350 disambiguates pronouns. In such implementations, the disambiguation module 350 identifies one of the one or more terms as a pronoun and determines the noun to which the pronoun refers. In some implementations, the disambiguation module 350 determines the noun to which the pronoun refers using a contact list associated with the user of the electronic device. Alternatively, or in addition, the disambiguation module 350 determines the noun to which the pronoun refers as the name of an entity, action, or location identified in a previous voice input associated with a previously tagged digital photograph. Alternatively, or in addition, the disambiguation module 350 determines the noun to which the pronoun refers as the name of a person identified based on a previous voice input associated with a previously tagged digital photograph. In some implementations, the disambiguation module 350 accesses information obtained from one or more sensors (e.g., proximity sensor 214, light sensor 212, GPS receiver 213, temperature sensor 215, motion sensor 210) of the handheld electronic device (e.g., user device 104) to determine the meaning of one or more of the terms. In some implementations, the disambiguation module 350 identifies two terms, each associated with either an entity, an action, or a location. For example, a first of the two terms refers to a person and a second of the two terms refers to a location. In some implementations, the disambiguation module 350 identifies three terms, each associated with either an entity, an action, or a location.
[0087] Once the task flow processor 336 completes the structured query for the actionable intent, the task flow processor 336 proceeds to execute the final task associated with the actionable intent. In response, the task flow processor 336 executes the steps and instructions in the task flow model according to the specific parameters included in the structured query. For example, a task flow model for the actionable intent of "restaurant reservation" may include steps and instructions for contacting a restaurant and requesting a reservation for a specific number of parties at a specific time. For example, using a structured query such as {restaurant reservation, restaurant=ABC Cafe, date=3 / 12 / 2012, time=7 PM, number of parties=5}, the task flow processor 336 may (1) log in to ABC Cafe's server or to a restaurant reservation system configured to accept reservations for multiple restaurants, such as ABC Cafe, (2) enter the date, time, and number of parties information into a form on the website, (3) submit the form, and (4) calendar the reservation into the user's calendar. In another example, described in more detail below, task flow processor 336 may cooperate with photo module 132 to perform steps and instructions associated with tagging or searching digital photos in response to voice input, for example. In some implementations, task flow processor 336 employs the assistance of service processing module 338 ("service processor") to complete tasks requested by user input or provide answers to information requested by user input. For example, service processor 338 may perform actions on behalf of task flow processor 336, such as placing phone calls, setting calendar entries, invoking map searches, invoking or interacting with other user applications installed on the user equipment, and invoking or interacting with third-party services (e.g., restaurant reservation portals, social networking websites or services, banking portals, etc.).In some embodiments, the protocols and application programming interfaces (APIs) required by each service may be specified by a respective service model in service models 356. Service processor 338 accesses the appropriate service model for the service and generates a request for the service according to the protocols and APIs required by the service associated with the service model.
[0088] For example, if a restaurant enables an online reservation service, the restaurant can present a service model that specifies the parameters needed to make a reservation and an API for communicating the values of the required parameters to the online reservation service. Upon request by task flow processor 336, service processor 338 can establish a network connection with the online reservation service using the web address stored in service model 356 and transmit the required reservation parameters (e.g., time, date, number of parties) to the online reservation interface in a format that complies with the online reservation service's API.
[0089] In some implementations, the natural language processor 332, dialog processor 334, and task flow processor 336 are used collaboratively and iteratively to infer and clarify the user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response that fulfills the user's intent (e.g., provides output to the user or completes a task).
[0090] In some embodiments, after all tasks necessary to fulfill the user's request have been performed, digital assistant 326 formulates an acknowledgment response and sends the response back to the user through I / O processing module 328. If the user request asks for an informational response, the acknowledgment response presents the requested information to the user. In some embodiments, the digital assistant also requests the user to indicate whether they are satisfied with the response created by digital assistant 326.
[0091] Attention is now directed to FIG. 4, a block diagram illustrating components of a voice trigger system 400, according to some embodiments. (The voice trigger system 400 is not limited to speech; the embodiments described herein equally apply to non-voice sounds.) The voice trigger system 400 comprises various components, modules, and / or software programs within the electronic device 104. In some embodiments, the voice trigger system 400 includes a noise detector 402, a sound type detector 404, a trigger sound detector 406, a speech-based service 408, and an audio subsystem 226, each coupled to an audio bus 401. In some embodiments, more or fewer modules are used. The sound detectors 402, 404, and 406 may be referred to as modules, and may include hardware (e.g., circuits, memory, processors, etc.), software (e.g., programs, software on a chip, firmware, etc.), and / or any combination thereof, for performing the functions described herein. In some embodiments, the sound detectors are communicatively, programmatically, physically, and / or operatively coupled to one another (e.g., via a communications bus), as indicated by the dashed lines in Figure 4. (For ease of explanation, Figure 4 shows each sound detector coupled only to adjacent sound detectors. It will be understood that each sound detector may be similarly coupled to any of the other sound detectors.)
[0092] In some implementations, the audio subsystem 226 includes a codec 410, an audio digital signal processor (DSP) 412, and a memory buffer 414. In some implementations, the audio subsystem 226 is coupled to one or more microphones 230 (FIG. 2) and one or more speakers 228 (FIG. 2). The audio subsystem 226 provides sound input to sound detectors 402, 404, 406 and speech-based services 408 (as well as other components or modules, such as a telephone and / or a telephone baseband subsystem) for processing and / or analysis. In some implementations, the audio subsystem 226 is coupled to an external audio system 416 that includes at least one microphone 418 and at least one speaker 420.
[0093] In some implementations, speech-based service 408 is a voice-based digital assistant and corresponds to one or more components or functions of the digital assistant system described above in connection with FIGS. 1-3C. In some implementations, the speech-based service is a speech-to-text service, a dictation service, etc. In some implementations, noise detector 402 monitors an audio channel to determine whether sound input from audio subsystem 226 meets a predetermined condition, such as an amplitude threshold. An audio channel corresponds to a stream of audio information received by one or more sound-receiving devices, such as one or more microphones 230 (FIG. 2). An audio channel may refer to audio information regardless of its processing state, or may refer to specific hardware processing and / or transmitting the audio information. For example, an audio channel may refer to analog electrical impulses from microphone 230 (and / or the circuitry through which they are propagated) as well as a digitally encoded audio stream resulting from processing of the analog electrical impulses (e.g., by audio subsystem 226 and / or any other audio processing system of electronic device 104).
[0094] In some embodiments, the predetermined condition is whether the sound input exceeds a particular volume for a predetermined time. In some embodiments, the noise detector uses a time-domain analysis of the sound input, which requires relatively few computational and battery resources compared to other types of analysis (e.g., as performed by sound type detector 404, trigger word detector 406, and / or speech-based service 408). In some embodiments, other types of signal processing and / or audio analysis are used, including, for example, frequency-domain analysis. When noise detector 402 determines that the sound input meets the predetermined condition, it activates an upstream sound detector, such as sound type detector 404 (e.g., by providing a control signal to initiate one or more processing routines and / or by providing power to the upstream sound detector). In some embodiments, the upstream sound detector is activated in response to another condition being met. For example, in some embodiments, the upstream sound detector is activated in response to determining that the device is not stored in an enclosed space (e.g., based on a light detector detecting a threshold level of light).
[0095] The sound type detector 404 monitors the audio channel and determines whether the sound input corresponds to a particular type of sound, such as a sound characteristic of a human voice, a whistle, applause, etc. The types of sounds that the sound type detector 404 is configured to recognize correspond to the particular trigger sound(s) that the voice trigger is configured to recognize. In embodiments where the trigger sound is a spoken word or phrase, the sound type detector 404 includes a "voice activity detector" (VAD). In some embodiments, the sound type detector 404 uses frequency domain analysis of the sound input. For example, the sound type detector 404 generates a spectrogram of the received sound input (e.g., using a Fourier transform) to analyze the spectral components of the sound input and determine whether the sound input appears to correspond to a particular type or category of sound (e.g., human speech). Thus, in embodiments where the trigger sound is a spoken word or phrase, if the audio channel picks up background sounds (e.g., traffic noise) rather than human speech, the VAD will not activate the trigger sound detector 406. In some implementations, sound type detector 404 remains enabled as long as a predetermined condition of any downstream sound detector (e.g., noise detector 402) is met. For example, in some implementations, sound type detector 404 remains enabled as long as the sound input (as determined by noise detector 402) includes sound above a predetermined amplitude threshold, and is disabled when the sound falls below the predetermined threshold. In some implementations, once activated, sound type detector 404 remains enabled until a condition is met, such as the expiration of a timer (e.g., 1, 2, 5, or 10 seconds, or any other suitable duration), the expiration of a particular number of on / off cycles of sound type detector 404, or the occurrence of an event (e.g., the amplitude of the sound falls below a second threshold, as determined by noise detector 402 and / or sound type detector 404).
[0096] As described above, when sound type detector 404 determines that the sound input corresponds to a predetermined type of sound, it activates an upstream sound detector, such as trigger sound detector 406 (e.g., by providing a control signal to initiate one or more processing routines and / or by providing power to the upstream sound detector).
[0097] The trigger sound detector 406 is configured to determine whether the sound input contains at least a portion of certain predetermined content (e.g., at least a portion of a trigger word, phrase, or sound). In some implementations, the trigger sound detector 406 compares a representation of the sound input (the "input representation") to one or more reference representations of the trigger word. If the input representation matches at least one of the one or more reference representations with an acceptable confidence value, the trigger sound detector 406 initiates the speech-based service 408 (e.g., by providing control signals to initiate one or more processing routines and / or by providing power to an upstream sound detector). In some implementations, the input representation and the one or more reference representations are spectrograms (or mathematical representations thereof), which describe how the spectral density of a signal changes over time. In some implementations, the representations are other types of audio signatures or voiceprints. In some embodiments, initiating the speech-based service 408 includes bringing one or more circuits, programs, and / or processors out of standby mode and invoking a sound-based service. The sound-based service then prepares to provide more comprehensive speech recognition, speech-to-text processing, and / or natural language processing. In some embodiments, the voice trigger system 400 includes voice authentication functionality to determine whether sound input corresponds to the voice of a particular person, such as the device owner / user. For example, in some embodiments, the sound type detector 404 uses voice printing techniques to determine whether the sound input was spoken by an authorized user. Voice authentication and voice printing are described in more detail in commonly-owned U.S. patent application Ser. No. 13 / 053,144, the entire contents of which are incorporated herein by reference. In some embodiments, voice authentication is included in any of the sound detectors described herein (e.g., noise detector 402, sound type detector 404, trigger sound detector 406, and / or speech-based service 408).In some implementations, voice authentication is implemented as a separate module from the sound detectors described above (e.g., as voice authentication module 428, FIG. 4), and may be operably located after noise detector 402, after sound type detector 404, after trigger sound detector 406, or in any other suitable location.
[0098] In some embodiments, trigger sound detector 406 remains enabled as long as the conditions of any downstream sound detector(s) (e.g., noise detector 402 and / or sound type detector 404) are met. For example, in some embodiments, trigger sound detector 406 remains enabled as long as the sound input includes a sound above a predetermined threshold (as detected by noise detector 402). In some embodiments, it remains enabled as long as the sound input includes a particular type of sound (as detected by sound type detector 404). In some embodiments, it remains enabled as long as both of the aforementioned conditions are met.
[0099] In some embodiments, once activated, trigger sound detector 406 remains enabled until a condition is met, such as the expiration of a timer (e.g., 1, 2, 5, or 10 seconds, or any other suitable duration), the expiration of a specific number of on / off cycles of trigger sound detector 406, or the occurrence of an event (e.g., the amplitude of the sound drops below a second threshold). In some embodiments, when one sound detector activates another, both sound detectors remain enabled. However, sound detectors may be enabled or disabled multiple times, and it is not necessary for all downstream (e.g., low-power and / or sophisticated) sound detectors to be enabled (or for each condition to be met) in order for an upstream sound detector to be enabled. For example, in some embodiments, after noise detector 402 and sound type detector 404 determine that each of their conditions have been met and trigger sound detector 406 is activated, one or both of noise detector 402 and sound type detector 404 are disabled and / or in standby mode while trigger sound detector 406 is operating. In other embodiments, both (or one or the other) of the noise detector 402 and sound type detector 404 remain enabled during operation of the trigger sound detector 406. In various embodiments, different combinations of sound detectors are enabled at different times, and whether one is enabled or disabled may depend on the state of the other sound detectors or may be independent of the state of the other sound detectors.
[0100] While Figure 4 illustrates three separate sound detectors, each configured to detect different types of sound input, various implementations of the voice trigger may use more or fewer sound detectors. For example, in some implementations, only trigger sound detector 406 is used. In some implementations, trigger sound detector 406 is used in conjunction with either noise detector 402 or sound type detector 404. In some implementations, all of detectors 402-406 are used. In some implementations, additional sound detectors are included as well.
[0101] Moreover, different combinations of sound detectors may be used at different times. For example, the particular combination of sound detectors and how they interact may depend on one or more conditions, such as the context or operating state of the device. As one specific example, when the device is plugged in (and therefore not solely dependent on battery power), trigger sound detector 406 is enabled while noise detector 402 and sound type detector 404 remain disabled. As another example, when the device is in a pocket or backpack, all sound detectors are disabled. Cascading sound detectors as described above, in which detectors requiring more power are invoked only when needed by detectors requiring less power, may provide a power-saving voice trigger function. As described above, further power savings are achieved by operating one or more of the sound detectors according to a duty cycle. For example, in some implementations, noise detector 402 operates according to a duty cycle to effectively provide continuous noise detection even when the noise detector is at least temporarily disabled. In some implementations, noise detector 402 is on for 10 milliseconds and off for 90 milliseconds. In some implementations, the noise detector 402 is on for 20 milliseconds and off for 500 milliseconds, although other on and off durations are possible.
[0102] In some implementations, if the noise detector 402 detects noise during its "on" interval, the noise detector 402 remains on and further processes and / or analyzes the sound input. For example, the noise detector 402 may be configured to activate an upstream sound detector upon detecting a sound above a predetermined amplitude for a predetermined period of time (e.g., 100 milliseconds). Thus, if the noise detector 402 detects a sound above a predetermined amplitude during its 10 millisecond "on" interval, it does not immediately enter an "off" interval. Instead, the noise detector 402 remains enabled and continues to process the sound input and determine whether it exceeds a threshold for the entire predetermined duration (e.g., 100 milliseconds).
[0103] In some embodiments, sound type detector 404 operates according to a duty cycle. In some embodiments, sound type detector 404 is on for 20 milliseconds and off for 100 milliseconds. Other on and off durations are possible. In some embodiments, sound type detector 404 can determine whether a sound input corresponds to a predetermined type of sound during the "on" interval of its duty cycle. Thus, if sound type detector 404 determines during its "on" interval that a sound is of a particular type, sound type detector 404 activates trigger sound detector 406 (or any other upstream sound detector). Alternatively, in some embodiments, if sound type detector 404 detects a sound that can correspond to a predetermined type during its "on" interval, the detector does not immediately enter an "off" interval. Instead, sound type detector 404 remains enabled and continues processing sound input to determine whether it corresponds to a predetermined type of sound. In some embodiments, once the sound detector determines that a predetermined type of sound has been detected, it activates the trigger sound detector 406, which further processes the sound input to determine whether a trigger sound has been detected. Like the noise detector 402 and the sound type detector 404, in some embodiments, the trigger sound detector 406 operates according to a duty cycle. In some embodiments, the trigger sound detector 406 is on for 50 milliseconds and off for 50 milliseconds. Other on and off durations are possible. If the trigger sound detector 406 detects during its "on" interval that a sound that may correspond to a trigger sound is present, the detector does not immediately enter its "off" interval. Instead, the trigger sound detector 406 remains enabled and continues to process the sound input to determine whether it contains a trigger sound. In some embodiments, once such a sound is detected, the trigger sound detector 406 remains enabled and processes the audio for a predetermined duration, such as 1, 2, 5, or 10 seconds, or any other suitable duration. In some embodiments, the duration is selected based on the length of the particular trigger word or sound it is configured to detect. For example, if the trigger phrase is "To SIRI," the trigger word detector will run for approximately two seconds to determine if the sound input contains the phrase.
[0104] In some embodiments, some of the sound detectors operate according to a duty cycle, while others operate continuously when enabled. For example, in some embodiments, only the first sound detector operates according to a duty cycle (e.g., noise detector 402 of FIG. 4 ), while upstream sound detectors operate continuously once activated. In some other embodiments, noise detector 402 and sound type detector 404 operate according to a duty cycle, while trigger sound detector 406 operates continuously. Whether a particular sound detector operates continuously or according to a duty cycle depends on one or more conditions, such as the context or operating state of the device. In some embodiments, if the device is connected to a power source and does not rely solely on battery power, all of the sound detectors operate continuously once activated. In other embodiments, noise detector 402 (or any of the sound detectors) operates according to a duty cycle when the device is in a pocket or backpack (e.g., as determined by sensor and / or microphone signals), but operates continuously if it is determined that the device may not be stored. In some implementations, whether a particular sound detector operates continuously or according to a duty cycle depends on the device's battery charge level. For example, noise detector 402 operates continuously when the battery charge is above 50% and according to a duty cycle when the battery charge is below 50%. In some implementations, the voice trigger includes noise, echo, and / or sound cancellation functionality (collectively referred to as noise cancellation). In some implementations, noise cancellation is performed by audio subsystem 226 (e.g., by audio DSP 412). Noise cancellation reduces or removes unwanted noise or sounds from the sound input before it is processed by the sound detector. In some implementations, the unwanted noise is background noise from the user's environment, such as a fan or clicking from a keyboard. In some implementations, the unwanted noise is any sound above or below a predetermined amplitude or frequency.For example, in some embodiments, sounds above the typical human vocal range (e.g., 3,000 Hz) are filtered out or removed from the signal. In some embodiments, multiple microphones (e.g., microphone 230) are used to help determine which components of the received sound should be reduced and / or removed. For example, in some embodiments, audio subsystem 226 uses beamforming techniques to identify sounds or portions of the sound input that originate from a single point in space (e.g., the user's mouth). Audio subsystem 226 then focuses on sounds that are received equally by all microphones (e.g., background sounds that are not coming from any particular direction) by removing them from the sound input.
[0105] In some implementations, DSP 412 is configured to cancel or remove from the sound input any sounds being output by the device on which the digital assistant is running. For example, if audio subsystem 226 is outputting music, radio, podcasts, voice output, or any other audio content (e.g., via speaker 228), DSP 412 removes any output sounds picked up by the microphone and included in the sound input. Thus, the sound input does not include (or at least includes less of) the output sound. Accordingly, the sound input provided to the sound detector is cleaner and more accurate triggering. Aspects of noise cancellation are described in further detail in commonly assigned U.S. Patent No. 7,272,224, the entire contents of which are incorporated herein by reference.
[0106] In some embodiments, different sound detectors require the sound input to be filtered and / or preprocessed in different ways. For example, in some embodiments, the noise detector 402 is configured to analyze time-domain audio signals between 60 and 20,000 Hz, while the sound type detector is configured to perform frequency-domain analysis of audio between 60 and 3,000 Hz. Thus, in some embodiments, the audio DSP 412 (and / or other audio DSPs in the device 104) preprocesses the received audio according to the respective needs of the sound detector. In some embodiments, the sound detectors, on the other hand, are configured to filter and / or preprocess audio from the audio subsystem 226 according to their specific needs. In such cases, the audio DSP 412 may still perform noise cancellation before providing the sound input to the sound detector. In some embodiments, the context of the electronic device is used to help determine whether and how a voice trigger is activated. For example, if the device is in a pocket, purse, or backpack, the user is unlikely to invoke a speech-based service, such as a voice-based digital assistant. Also, a user is unlikely to invoke a speech-based service during a loud rock concert. Some users are unlikely to invoke a speech-based service at certain times (e.g., late at night). However, there are contexts in which a user may well invoke a speech-based service using a voice trigger. For example, some users may well use a voice trigger while driving, when alone, at work, etc. Various techniques are used to determine the context of a device. In various implementations, the device uses information from any one or more of the following components or sources to determine the context of the device: a GPS receiver, a light sensor, a microphone, a proximity sensor, an orientation sensor, an inertial sensor, a camera, communication circuitry and / or antenna, charging circuitry and / or power circuitry, switch position, a temperature sensor, a compass, an accelerometer, a calendar, user preferences, etc.The device context can then be used to adjust whether and how the voice trigger operates. For example, in certain contexts, the voice trigger may be disabled (or operated in a different mode) as long as the context is maintained. For example, in some implementations, the voice trigger may be disabled when the phone is in a certain orientation (e.g., placed face down on a surface), during a certain period of time (e.g., between 10:00 PM and 8:00 AM), when the phone is in "silent" or "do not disturb" mode (e.g., based on a switch position, mode setting, or user preference), when the device is in a substantially enclosed space (e.g., a pocket, bag, purse, drawer, or glove box), when the device is near other devices that have voice triggers and / or speech-based services (e.g., based on proximity sensors, voice / wireless / infrared communications), etc. In some embodiments, instead of being disabled, the voice trigger system 400 is operated in a low power mode (e.g., by operating the noise detector 402 according to a duty cycle with a 10 millisecond "on" interval and a 5 second "off" interval). In some embodiments, the audio channel is monitored less frequently when the voice trigger system 400 is operated in a low power mode. In some embodiments, the voice trigger uses a different sound detector or combination of sound detectors when in a low power mode than when in a normal mode. (The voice trigger may be capable of many different modes or operating states, each of which may use different amounts of power, and different embodiments use them according to their particular designs.)
[0107] On the other hand, if the device is in some other context, the voice trigger remains active (or operates in a different mode) as long as the context is maintained. For example, in some implementations, the voice trigger remains active while the phone is connected to a power source, while the phone is in a predetermined orientation (e.g., placed face up on a surface), during a predetermined period of time (e.g., between 8:00 AM and 10:00 PM), while the device is moving and / or in a vehicle (e.g., based on a GPS signal, a BLUETOOTH connection, or while connected to a vehicle, etc.). Aspects of detecting evidence that a device is in a vehicle are described in more detail in commonly assigned U.S. Provisional Patent Application No. 61 / 657,744, the entire contents of which are incorporated herein by reference. Various examples of methods for determining a particular context are provided below. In various embodiments, these and other contexts are detected using different techniques and / or information sources.
[0108] As described above, whether the voice trigger system 400 is enabled (e.g., listening) may depend on the physical orientation of the device. In some implementations, the voice trigger is enabled when the device is placed “face up” on a surface (e.g., with the display and / or touchscreen surface visible) and / or disabled when placed “face down.” This provides a user with an easy way to enable and / or disable the voice trigger without having to navigate settings menus, switches, or buttons. In some implementations, the device detects whether it is placed face up or face down on a surface using a light sensor (e.g., based on the difference in incident light on the front and back of the device 104), a proximity sensor, a magnetic sensor, an accelerometer, a gyroscope, a tilt sensor, a camera, etc. In some implementations, other operating modes, settings, parameters, or preferences are affected by the orientation and / or position of the device. In some implementations, the particular trigger sound, word, or phrase that the voice trigger is listening for depends on the orientation and / or position of the device. For example, in some implementations, the voice trigger listens for a first trigger word, phrase, or sound when the device is in one orientation (e.g., face-up on a surface) and a different trigger word, phrase, or sound when the device is in another orientation (e.g., face-down). In some implementations, the trigger phrase for the face-down orientation is longer and / or more complex than that for the face-up orientation. Thus, a user can place the device face-down when other people are around or in a noisy environment and still activate the voice trigger while also reducing fraudulent acceptances that would be more frequent for shorter or simpler trigger words. As one specific example, the face-up trigger phrase may be "To SIRI," while the face-down trigger phrase may be "To SIRI, this is Andrew, please wake me up." Longer trigger phrases also provide the sound detector and / or voice authenticator with longer voice samples to process and / or analyze, thus increasing the accuracy of the voice trigger and reducing fraudulent acceptances.
[0109] In some embodiments, the device 104 detects whether it is in a vehicle (e.g., an automobile). Voice triggers are particularly useful for invoking speech-based services when the user is in a vehicle because they help reduce the physical interaction required to operate the device and / or speech-based services. Indeed, one advantage of voice-based digital assistants is that they can be used to perform tasks when it is impossible or dangerous to see and touch the device. Thus, voice triggers may be used when the device is in a vehicle so that the user does not need to touch the device to invoke the digital assistant. In some embodiments, the device determines whether it is in a vehicle by detecting that it is connected to and / or paired with the vehicle, such as through BLUETOOTH communication (or other wireless communication) or a docking connector or cable. In some embodiments, the device determines whether it is in a vehicle by determining the device's location and / or velocity (e.g., using a GPS receiver, accelerometer, and / or gyroscope). For example, if the device is determined to be traveling above 20 mph and along a road, and therefore is likely to be inside a vehicle, the voice trigger may continue to remain enabled and / or in a high power or high sensitivity state.
[0110] In some embodiments, the device detects whether the device is stored (e.g., in a pocket, purse, bag, drawer, etc.) by determining whether it is within a substantially enclosed space. In some embodiments, the device determines whether it is stored using a light sensor (e.g., a dedicated ambient light sensor and / or a camera). For example, in some embodiments, the device is likely stored if the light sensor detects low or no light. In some embodiments, the time of day and / or the device's location are also taken into consideration. For example, if a light sensor detects low light levels when high light levels are expected (e.g., daytime), the device may be stored and voice trigger system 400 may not be needed. Thus, voice trigger system 400 may enter a low-power or standby state. In some embodiments, differences in light detected by sensors located on opposite sides of the device may be used to determine its location, and therefore whether it is stored. Specifically, if the device is not stored in a pocket or bag but is placed on a table or surface, the user may attempt to activate the voice trigger. When the device is placed face down (or face up) on a surface such as a table or desk, one side of the device is covered, exposing the other side to ambient light while the other side is exposed to low or no light. Thus, if the light sensors on the front and back of the device detect significantly different light levels, the device is determined to be unstored. On the other hand, if the light sensors on the opposite side detect the same or similar light levels, the device is determined to be stored in a substantially enclosed space. Also, if both light sensors detect low light levels during the day (or if the device expects the phone to be in a brightly lit environment), the device is determined to be stored with high confidence.
[0111] In some embodiments, other techniques are used (instead of or in addition to optical sensors) to determine whether the device is stored. For example, in some embodiments, the device emits one or more sounds (e.g., tones, clicks, pings, etc.) from a speaker or transducer (e.g., speaker 228) and monitors one or more microphones or transducers (e.g., microphone 230) to detect echoes of the omission sound(s). (In some embodiments, the device emits inaudible signals, such as sounds outside the range of human hearing.) From the echoes, the device determines characteristics of the surrounding environment. For example, a relatively large environment (e.g., a room or car interior) reflects sound differently than a relatively small, enclosed environment (e.g., a pocket, purse, bag, drawer, etc.).
[0112] In some implementations, the voice trigger system 400 operates differently when it is near other devices (such as other devices with voice triggers and / or speech-based services) than when it is far away from the other devices. This can be useful, for example, when many devices are close to each other, to disable or desensitize the voice trigger system 400 so that when one person speaks a trigger word, other surrounding devices are not similarly triggered. In some implementations, a device determines its proximity to other devices using RFID, proximity communication, infrared / acoustic signals, etc. As described above, voice triggers are particularly useful when a device is operated in a hands-free mode, such as when a user is driving. In such cases, users often use external audio systems, such as wired or wireless headsets, watches with speakers and / or microphones, vehicle-integrated microphones and speakers, etc., to make calls or dictate text input without having to hold the device close to their face. For example, wireless headsets and vehicle audio systems may connect to electronic devices using BLUETOOTH® communication or any other suitable wireless communication. However, this can be inefficient for a voice trigger that monitors audio received via a wireless audio accessory due to the power required to maintain an open audio channel with the wireless accessory. Wireless headsets, in particular, can retain enough power in their batteries to provide several hours of continuous talk time, and are therefore well-suited to conserving the battery for when the headset is needed for actual communication, instead of using it simply to monitor ambient audio and wait for potential trigger sounds. Furthermore, wired external headset accessories may require more power than an on-board microphone alone, and keeping the headset's microphone active drains the device's battery charge. This is especially true considering that the ambient audio received by a wireless or wired headset typically consists largely of silence or irrelevant sounds.Thus, in some implementations, the voice trigger system 400 monitors audio from the on-device microphone 230, even if the device is coupled to an external microphone (wired or wireless). If the voice trigger then detects a trigger word, the device initiates an active audio link with the external microphone and receives subsequent sound input (such as a command to a voice-based digital assistant) via the external microphone rather than the on-device microphone 230. Provided certain conditions are met, an active communication link may be maintained between the device and an external audio system 416 (which may be communicatively coupled to the device 104 via wired or wireless), and the voice trigger system 400 may listen for the trigger sound via the external audio system 416 instead of (or in addition to) the on-device microphone 230. For example, in some implementations, movement characteristics of the electronic device and / or external audio system 416 (e.g., determined by an accelerometer, gyroscope, etc. on each device) are used to determine whether the voice trigger system 400 should monitor background sounds using the on-device microphone 230 or the external microphone 418. Specifically, the difference in movement between the device and the external audio system 416 provides information about whether the external audio system 416 is actually in use. For example, if both the device and the wireless headset are moving (or not moving) substantially equally, it may be determined that the headset is not in use or not being worn. This may occur, for example, because both devices are close to each other and idle (e.g., resting on a table or in a pocket, bag, purse, drawer, etc.). Accordingly, under these conditions, the voice trigger system 400 monitors the on-device microphone because it is unlikely that the headset is actually in use. If there is a difference in movement between the wireless headset and the device, it may be determined that the user is wearing the headset.These conditions may occur, for example, while the headset is being worn on a user's head because the device is placed (e.g., on a surface or in a bag) (where at least a small amount of movement is likely even if the wearer is relatively stationary). Under these conditions, the headset is considered to be worn, so the voice trigger system 400 maintains an active communications link and monitors the headset's microphone 418 instead of (or in addition to) the microphone 230 on the device. This technique focuses on differences in the movement of the device and the headset, so that movement common to both devices is canceled out. This may be useful, for example, when a user is using the headset in a moving vehicle, where the device (e.g., a cell phone) is in a cup holder, on an empty seat, or in the user's pocket, and the headset is worn on the user's head. Once movement common to both devices is canceled out (e.g., vehicle movement), the relative movement (if any) of the headset compared to the device can be determined to determine whether the headset is likely in use (or whether the headset is not being worn). Although the above description refers to wireless headsets, similar techniques apply to wired headsets as well.
[0113] Because human voices vary widely, it may be necessary or useful to tune a voice trigger to improve its accuracy in recognizing a particular user's voice. Additionally, a person's voice may change over time due to, for example, natural voice changes due to illness, aging, or hormonal changes. Accordingly, in some implementations, the voice trigger system 400 can adapt its voice and / or sound recognition profile to a particular user or group of users. As described above, a sound detector (e.g., sound type detector 404 and / or trigger sound detector 406) may be configured to compare a representation of a sound input (e.g., a sound or utterance provided by a user) to one or more reference representations. For example, if the input representation matches the reference representations with a predetermined confidence level, the sound detector determines that the sound input corresponds to a predetermined type of sound (e.g., sound type detector 404) or that the sound input contains predetermined content (e.g., trigger sound detector 406). To tune the voice trigger system 400, in some embodiments, the device prepares a reference representation to which the input representation is compared. In some embodiments, the reference representation is prepared (or created) as part of a voice enrollment or "training" procedure, in which the user outputs a trigger sound several times to allow the device to tune (or create) the reference representation. The device then creates the reference representation using the person's actual voice.
[0114] In some embodiments, the device uses trigger sounds received under normal use conditions to adjust the reference representation. After a successful voice triggering event (e.g., a sound input is found that meets all of the triggering criteria), for example, the device uses information from the sound input to adjust and / or tune the reference representation. In some embodiments, only sound inputs that are determined to meet all or part of the triggering criteria with a certain confidence level are used to adjust the reference representation. Thus, if the voice trigger has low confidence that a sound input corresponds to or includes a trigger sound, that sound input may be ignored for purposes of adjusting the reference representation. However, in some embodiments, sound inputs that meet the voice trigger system 400 with low confidence are used to adjust the reference representation.
[0115] In some embodiments, the device 104 iteratively adjusts the reference representation (using these or other techniques) as more and more sound input is received to accommodate subtle changes in the user's voice over time. For example, in some embodiments, the device 104 (and / or associated devices or services) adjusts the reference representation after each successful triggering event. In some embodiments, the device 104 analyzes the sound input associated with each successful triggering event, determines whether the reference representation should be adjusted based on that input (e.g., if certain conditions are met), and adjusts the reference representation only if it is appropriate to do so. In some embodiments, the device 104 maintains a running average of the reference representation over time. In some embodiments, the voice trigger system 400 detects sounds that do not meet one or more of the triggering criteria (e.g., as determined by one or more of the sound detectors), which may be an actual attempt by a legitimate user to do so. For example, the voice trigger system 400 may be configured to respond to a trigger phrase such as "Siri," but if the user's voice changes (e.g., due to illness, aging, changes in accent / tone, etc.), the voice trigger system 400 may not recognize the user's attempt to activate the device. (This may also occur if the voice trigger system 400 is set to a default condition and / or is not properly tuned to the user's particular voice, such as if the user has not initialized or gone through a training procedure to customize the voice trigger system 400 for the user's voice.) If the voice trigger system 400 does not respond to the user's first attempt to activate the voice trigger, perhaps the user repeats the trigger phrase. The device detects that these repeated sound inputs are similar to each other and / or to the trigger phrase (even if they are not similar enough to cause the voice trigger system 400 to activate the speech-based service).If such conditions are met, the device determines that the sound inputs correspond to a legitimate attempt to activate the voice trigger system 400. In response, in some embodiments, the voice trigger system 400 uses the received sound inputs to adjust one or more aspects of the voice trigger system 400 so that similar utterances by the user are recognized as legitimate triggers in the future. In some embodiments, these sound inputs are used to adapt the voice trigger system 400 only if a specific condition or combination of conditions is met. For example, in some embodiments, the sound inputs are used to adapt the voice trigger system 400 when a predetermined number of sound inputs are received consecutively (e.g., 2, 3, 4, 5, or any other suitable number), when the sound inputs are sufficiently similar to a reference representation, when the sound inputs are sufficiently similar to each other, when the sound inputs are close to each other (e.g., received within a predetermined period and / or at or near a predetermined interval), and / or under any combination of these or other conditions. In some cases, the voice trigger system 400 may detect one or more sound inputs that do not meet one or more of the triggering criteria followed by a manual initiation of a speech-based service (e.g., by pressing a button or icon). In some implementations, the voice trigger system 400 determines that the sound inputs indeed correspond to a failed voice triggering attempt because the speech-based service was initiated shortly after receiving the sound inputs. In response, the voice trigger system 400 uses those received sound inputs, as described above, to adjust one or more aspects of the voice trigger system 400 so that user utterances are recognized as legitimate triggers in the future.
[0116] While the adaptation techniques described above refer to adjusting the reference representation, other aspects of the trigger sound detection technique may be adjusted in the same or similar manner in addition to or instead of adjusting the reference representation. For example, in some embodiments, the device adjusts how the sound input is filtered and / or what filters are applied to the sound input, such as focusing on and / or reducing specific frequencies or frequency ranges of the sound input. In some embodiments, the device adjusts the algorithm used to compare the input representation to the reference representation. For example, in some embodiments, one or more terms of a mathematical function used to determine the difference between the input representation and the reference representation are changed, added, or removed, or replaced with a different mathematical function. In some embodiments, adaptation techniques such as those described above require more resources than the voice trigger system 400 is able or configured to provide. In particular, the sound detector may not have the amount or type of, or access to, the processor, data, or memory necessary to perform iterative adaptation of the reference representation and / or sound detection algorithm (or any other suitable aspect of the voice trigger system 400). Accordingly, in some embodiments, one or more of the above adaptation techniques are performed by a more powerful processor, such as an application processor (e.g., processor(s) 204), or by a different device (e.g., the server system 108). However, the voice trigger system 400 is designed to operate even when the application processor is in standby mode. Thus, the sound input used to adapt the voice trigger system 400 is received when the application processor is not available and cannot process the sound input. Accordingly, in some embodiments, the sound input is stored by the device so that it can be further processed and / or analyzed after receipt. In some embodiments, the sound input is stored in a memory buffer 414 of the audio subsystem 226.In some embodiments, the sound input is stored in system memory (e.g., memory 250, FIG. 2) using direct memory access (DMA) techniques (e.g., including using a DMA engine to copy or move data without having to wake up the application processor). The stored sound input is then provided to or accessed by the application processor (or server system 108, or another suitable device) so that, upon wake-up, the application processor can perform one or more of the adaptation techniques described above. In some embodiments,
[0117] 5-7 are flow diagrams illustrating a method for operating a voice trigger according to certain embodiments. The method is optionally governed by instructions stored in computer memory or a persistent computer-readable storage medium (e.g., memory 250 of client device 104, memory 302 associated with digital assistant system 300) and executed by one or more processors of one or more computer systems of the digital assistant system, including, but not limited to, server system 108 and / or user device 104a. The computer-readable storage medium may include magnetic or optical disk storage, solid-state storage such as flash memory, or other non-volatile memory device(s). The computer-readable instructions stored on the computer-readable storage medium may include one or more of source code, assembly language code, object code, or other instruction formats interpreted and executed by one or more processors. In various embodiments, some operations of the respective methods shown in the figures may be combined and / or the order of some operations may be changed from the order shown. Also, in some implementations, operations shown in separate figures and / or described in connection with separate methods may be combined to form other methods, and operations described in connection with the same figure and / or the same method may be separated into different methods. Moreover, in some implementations, one or more operations in a method are performed by modules of the digital assistant system 300 and / or electronic device (e.g., user device 104), including, for example, the natural language processing module 332, the dialog flow processing module 334, the audio subsystem 226, the noise detector 402, the sound type detector 404, the trigger sound detector 406, the speech-based service 408, and / or any submodules thereof. FIG. 5 illustrates a method 500 for operating a voice trigger system (e.g., the voice trigger system 400 of FIG. 4, FIG. 4) according to some implementations. In some implementations, the method 500 is performed in an electronic device including one or more processors and a memory that stores instructions executed by the one or more processors (e.g., the electronic device 104).The electronics receives audio input (502). The audio input may correspond to speech (e.g., a word, phrase, or sentence), human pronunciation (e.g., whistling, clicking, snapping, clapping, etc.), or any other sound (e.g., electronically generated chirps, mechanical noise makers, etc.). In some implementations, the electronics receives audio input via audio subsystem 226 (e.g., including codec 410, audio DSP 412, and buffer 414, as well as microphones 230 and 418, described in connection with FIG. 4).
[0118] In some embodiments, the electronics determines whether the sound input meets a predetermined condition (504). In some embodiments, the electronics applies time-domain analysis to the sound input to determine whether the sound input meets a predetermined condition. For example, the electronics analyzes the sound input over a period of time to determine whether the sound amplitude reaches a predetermined level. In some embodiments, a threshold is met when the amplitude (e.g., volume) of the sound input meets and / or exceeds a predetermined threshold. In some embodiments, it is met when the sound input meets and / or exceeds a predetermined threshold for a predetermined time. As described in more detail below, in some embodiments, determining whether the sound input meets a predetermined condition (504) is performed by a third sound detector (e.g., noise detector 402). (The term third sound detector is used in this case to distinguish this sound detector from the other sound detectors (e.g., the first and second sound detectors described below) and does not necessarily indicate any operational position or order of the sound detectors.)
[0119] The electronic device determines whether the sound input corresponds to a predetermined type of sound (506). As described above, sounds are classified into various "types" based on certain distinguishable sound characteristics. Determining whether the sound input corresponds to a predetermined type includes determining whether the sound input includes or exhibits characteristics of the particular type. In some embodiments, the predetermined type of sound is a human voice. In such embodiments, determining whether the sound input corresponds to a human voice includes determining whether the sound input includes a frequency characteristic of a human voice (508). As described in more detail below, in some embodiments, determining whether the sound input corresponds to a predetermined type of sound (506) is performed by a first sound detector (e.g., sound type detector 404). Upon determining that the sound input corresponds to a predetermined type of sound, the electronic device determines whether the sound input includes predetermined content (510). In some embodiments, the predetermined content corresponds to one or more predetermined phonemes (512). In some embodiments, the one or more predetermined phonemes comprise at least one word. In some implementations, the predetermined content is a sound (e.g., a whistle, a click, or a clap). In some implementations, determining 510 whether the sound input includes the predetermined content is performed by a second sound detector (e.g., trigger sound detector 406), as described below.
[0120] Upon determining that the audio input includes predetermined content, the electronic device initiates (514) the speech-based service. In some embodiments, the speech-based service is a voice-based digital assistant, as described in detail above. In some embodiments, the speech-based service is a dictation service, where speech input is converted to text and included and / or displayed in a text entry field (e.g., in an email, text message, word processing, or note-taking application, etc.). In embodiments where the speech-based service is a voice-based digital assistant, when the voice-based digital assistant is initiated, a prompt (e.g., a sound or speech prompt) is issued to the user indicating that the user can provide audio input and / or commands to the digital assistant. In some embodiments, initiating the voice-based digital assistant includes enabling an application processor (e.g., processor(s) 204, FIG. 2), initiating one or more programs or modules (e.g., digital assistant client module 264, FIG. 2), and / or establishing a connection to a remote server or device (e.g., digital assistant server 106, FIG. 1).
[0121] In some embodiments, the electronic device determines whether the sound input corresponds to the voice of the particular user (516). For example, one or more voice authentication techniques are applied to the sound input to determine whether it corresponds to the voice of an authorized user of the device. Voice authentication techniques are described in detail above. In some embodiments, the voice authentication is performed by one of the sound detectors (e.g., trigger sound detector 406). In some embodiments, the voice authentication is performed by a dedicated voice authentication module (including any suitable hardware and / or software). In some embodiments, a sound-based service is initiated in response to determining that the sound input contains predetermined content and corresponds to the voice of the particular user. Thus, for example, a sound-based service (e.g., a voice-based digital assistant) is initiated only when a trigger word or phrase is spoken by an authorized user. This reduces the likelihood that the service can be invoked by unauthorized users and can be particularly useful when multiple electronic devices are in close proximity, so that one user's utterance of a trigger sound does not activate another user's voice trigger.
[0122] In some implementations where the speech-based service is a voice-based digital assistant, in response to determining that the sound input includes predetermined content but does not correspond to a particular user's voice, the voice-based digital assistant is initiated in a restricted access mode. In some implementations, the restricted access mode allows the digital assistant to access only a subset of the data, services, and / or functionality that the digital assistant might otherwise provide. In some implementations, the restricted access mode corresponds to a write-only mode (e.g., so that unauthorized users of the digital assistant cannot access data from the calendar, task list, contacts, photos, email, text messages, etc.). In some implementations, the restricted access mode corresponds to a sandboxed instance of the speech-based service, preventing the speech-based service from reading from or writing to a user's data, such as user data 266 on device 104 (FIG. 2) or any other device (e.g., user data 348 of FIG. 3A, which may be stored on a remote server, such as server system 108 of FIG. 1).
[0123] In some embodiments, in response to determining that the sound input includes the predetermined content and that the sound input corresponds to the voice of the particular user, the voice-based digital assistant outputs a prompt that includes the name of the particular user. For example, once the particular user is identified through voice authentication, the voice-based digital assistant may output a prompt such as, "Peter, what can I do for you?" instead of a more general prompt, such as a tone, beep, or non-proprietary voice prompt. As described above, in some embodiments, a first sound detector determines whether the sound input corresponds to a predetermined type of sound (at step 506), and a second sound detector determines whether the sound detector includes the predetermined content (at step 510). In some embodiments, the first sound detector consumes less power during operation than the second sound detector, e.g., because the first sound detector uses less processor-intensive technology than the second sound detector. In some embodiments, the first sound detector is sound type detector 404, and the second sound detector is trigger sound detector 406, both of which are described above in connection with FIG. 4. In some implementations, during these operations, the first sound detector and / or the second sound detector periodically monitor the audio channel according to a duty cycle, as described above in connection with FIG.
[0124] In some embodiments, the first sound detector and / or the sound detectors perform a frequency domain analysis of the sound input. For example, these sound detectors perform a Laplace transform, a Z transform, or a Fourier transform to generate a frequency spectrum or determine the spectral density of the sound input or a portion thereof. In some embodiments, the first sound detector is a voice activity detector configured to determine whether the sound input includes frequencies characteristic of a human voice (or other features, aspects, or aspects of the sound input that are characteristic of a human voice).
[0125] In some implementations, the second sound detector is off or disabled until the first sound detector detects a predetermined type of sound input. Accordingly, in some implementations, method 500 includes activating the second sound detector in response to determining that the sound input corresponds to the predetermined type. (In other implementations, the second sound detector is activated in response to other conditions or is continuously operated regardless of the determination from the first sound detector.) In some implementations, activating the second sound detector includes enabling hardware and / or software (e.g., including circuits, processors, programs, memory, etc.). In some implementations, the second sound detector is operated (e.g., enabled and monitoring an audio channel) for at least a predetermined time period after activation. For example, when the first sound detector determines that the sound input corresponds to a predetermined type (e.g., includes a human voice), the second sound detector is activated and determines whether the sound input also includes predetermined content (e.g., a trigger word). In some implementations, the predetermined time period corresponds to the duration of the predetermined content. Thus, if the predetermined content is the phrase "Dear SIRI," the predetermined time will be long enough to determine whether the phrase was uttered (e.g., 1 or 2 seconds, or any other suitable duration). If the predetermined content is longer, such as the phrase "Dear SIRI, wake me up and help me," the predetermined time will be longer (e.g., 5 seconds, or another suitable duration). In some embodiments, the second sound detector is activated as long as the first sound detector detects a sound corresponding to the predetermined type. In such embodiments, for example, as long as the first sound detector detects a human voice in the sound input, the second sound detector processes the sound input and determines whether it contains the predetermined content.
[0126] As described above, in some embodiments, a third sound detector (e.g., noise detector 402) determines whether the sound input meets a predetermined condition (at step 504). In some embodiments, the third sound detector consumes less power during operation than the first sound detector. In some embodiments, the third sound detector periodically monitors the audio channel according to a duty cycle, as described above with respect to FIG. 4. Also, in some embodiments, the third sound detector performs a time-domain analysis of the sound input. In some embodiments, the third sound detector consumes less power than the first sound detector because the time-domain analysis is less processor-intensive than the frequency-domain analysis applied by the second sound detector.
[0127] Similar to the above discussion regarding activating a second sound detector (e.g., trigger sound detector 406) in response to a determination by a first sound detector (e.g., sound type detector 404), in some embodiments, the first sound detector is activated in response to a determination by a third sound detector (e.g., noise detector 402). For example, in some embodiments, sound type detector 404 is activated in response to a determination by noise detector 402 that the sound input meets a predetermined condition (e.g., exceeds a particular volume for a sufficient duration). In some embodiments, activating the first sound detector includes enabling hardware and / or software (e.g., including circuits, processors, programs, memory, etc.). In other embodiments, the first sound detector is activated in response to other conditions or is continuously operated. In some embodiments, the device stores (518) at least a portion of the sound input in memory. In some embodiments, the memory is buffer 414 of audio subsystem 226 (FIG. 4). The stored sound input allows the device to process the sound input in non-real-time. For example, in some embodiments, one or more of the sound detectors read and / or receive stored sound input and process the stored sound input. This may be particularly useful if an upstream sound detector (e.g., trigger sound detector 406) is not activated until partway through receipt of sound input by audio subsystem 226. In some embodiments, the stored portion of the sound input is provided (520) to a speech-based service when the speech-based service is initiated. Thus, the speech-based service can copy, process, or otherwise operate on the stored portion of the sound input, even though the speech-based service is not fully operational until the portion of the sound input is received. In some embodiments, the stored portion of the sound input is provided to an accommodation module of the electronic device.
[0128] In various embodiments, steps 516-520 occur at different locations within method 500. For example, in some embodiments, one or more of steps 516-520 occur between steps 502 and 504, between steps 510 and 514, or at any other suitable location.
[0129] FIG. 6 illustrates a method 600 for operating a voice trigger system (e.g., voice trigger system 400 of FIG. 4, FIG. 4) according to some embodiments. In some embodiments, method 600 is performed in an electronic device including one or more processors and a memory storing instructions executed by the one or more processors (e.g., electronic device 104). The electronic device determines whether it is in a predetermined orientation (602). In some embodiments, the electronic device detects its orientation using a light sensor (including a camera), a microphone, a proximity sensor, a magnetic sensor, an accelerometer, a gyroscope, a tilt sensor, etc. For example, the electronic device determines whether it is placed face-up or face-down on a surface by comparing the amount or brightness of light incident on a front-facing camera sensor with the amount or brightness of light incident on a rear-facing camera sensor. If the amount and / or brightness detected by the front-facing camera is significantly greater than that detected by the rear-facing camera, the electronic device is determined to be face-up. On the other hand, if the amount and / or brightness detected by the rear-facing camera is significantly greater than that detected by the front-facing camera, the electronic device is determined to be face-down. Upon determining that the electronic device is in a predetermined orientation, the electronic device enables a predetermined mode of the voice trigger (604). In some embodiments, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing downward, and the predetermined mode is a standby mode (606). For example, in some embodiments, when the smartphone or tablet is placed on a table or desk with the screen facing downward, the voice trigger goes into standby mode (e.g., powered off) to prevent unintentional activation of the voice trigger.
[0130] However, in some implementations, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing up, and the predetermined mode is listening mode 608. Thus, for example, if the smartphone or tablet is placed on a table or desk with the screen facing up, the voice trigger will be in listening mode and can respond to the user upon detecting the trigger.
[0131] 7 illustrates a method 700 for operating a voice trigger (e.g., voice trigger system 400, FIG. 4) according to some embodiments. In some embodiments, method 700 is performed on an electronic device including one or more processors and a memory that stores instructions executed by the one or more processors (e.g., electronic device 104). The electronic device operates (702) a voice trigger (e.g., voice trigger system 400) in a first mode. In some embodiments, the first mode is a normal listening mode.
[0132] The electronic device determines whether it is in a substantially enclosed space by detecting that one or more of the electronic device's microphone and camera are occluded 704. In some embodiments, the substantially enclosed space includes a pocket, purse, bag, drawer, glove box, briefcase, etc.
[0133] As described above, in some embodiments, the device detects that a microphone is occluded by emitting one or more sounds (e.g., tones, clicks, pings, etc.) from a speaker or transducer and monitoring one or more microphones or transducers to detect echoes of the occluded sound(s). For example, a relatively large environment (e.g., a room or car interior) reflects sound differently than a relatively small, substantially enclosed environment (e.g., a purse or pocket). Thus, when the device detects that a microphone (or the speaker that emitted the sound) is occluded based on the echo (or lack of echo), the device determines that it is in a substantially enclosed space. In some embodiments, the device detects that a microphone is occluded by detecting that the microphone picks up sounds characteristic of enclosed spaces. For example, if the device is in a pocket, the microphone may detect a characteristic soft noise caused by the microphone contacting or being in close proximity to the fabric of the pocket. In some implementations, the device detects that the camera is occluded based on the level of light received by the sensor or by determining whether it can obtain an in-focus image. For example, if the camera sensor detects a low level of light during a time when high levels of light are expected (e.g., daytime), the device determines that the camera is occluded and that the device is in a substantially enclosed space. As another example, the camera may attempt to obtain an in-focus image on its sensor. Typically, this is difficult when the camera is in a very dark location (e.g., a pocket or backpack) or is too close to the subject it is attempting to focus on (e.g., inside a purse or backpack). Thus, if the camera is unable to obtain an in-focus image, it determines that the device is in a substantially enclosed space.
[0134] Upon determining that the electronic device is within a substantially enclosed space, the electronic device switches the voice trigger to a second mode (706). In some embodiments, the second mode is a standby mode (708). In some embodiments, when in standby mode, the voice trigger system 400 continues to monitor ambient sounds but does not respond to received sounds, regardless of whether the voice trigger system 400 is otherwise activated. In some embodiments, in standby mode, the voice trigger system 400 is disabled and does not process sounds to detect trigger sounds. In some embodiments, the second mode includes operating one or more sound detectors of the voice trigger system 400 according to a different duty cycle than the first mode. In some embodiments, the second mode includes operating a different combination of sound detectors than the first mode.
[0135] In some embodiments, the second mode corresponds to a more sensitive monitoring mode, allowing voice trigger system 400 to detect and respond to trigger sounds even when within the substantially enclosed space. In some embodiments, once the voice trigger switches to the second mode, the device periodically determines whether the electronic device is still within the substantially enclosed space by detecting whether one or more of the electronic device's microphone and camera are occluded (e.g., using any of the techniques described above with respect to step (704)). If the device is still within the substantially enclosed space, voice trigger system 400 remains in the second mode. In some embodiments, once the device is removed from the substantially enclosed space, the electronic device switches the voice trigger back to the first mode.
[0136] According to some implementations, FIG. 8 illustrates a functional block diagram of an electronic device 800 configured in accordance with the principles of the present invention as described above. The functional blocks of this device may be implemented by hardware, software, or a combination of hardware and software to implement the principles of the present invention. It will be understood by those skilled in the art that the functional blocks described in FIG. 8 may be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description herein may support any possible combination or division, or definition of additional functional blocks described herein.
[0137] 8 , electronic device 800 includes a sound receiving unit 802 configured to receive sound input. Electronic device 800 also includes a processing unit 806 coupled to speech receiving unit 802. In some implementations, processing unit 806 includes a noise detector 808, a sound type detector 810, a trigger sound detector 812, a service initiation unit 814, and a voice authentication unit 816. In some implementations, noise detector 808 corresponds to noise detector 402 described above and is configured to perform any of the operations described above for noise detector 402. In some implementations, sound type detector 810 corresponds to sound type detector 404 described above and is configured to perform any of the operations described above for sound type detector 404. In some implementations, trigger sound detector 812 corresponds to trigger sound detector 406 described above and is configured to perform any of the operations described above for trigger sound detector 406. In some implementations, voice authentication unit 816 corresponds to voice authentication module 428 described above and is configured to perform any of the operations described above for voice authentication module 428. The processing unit 806 is configured to determine whether at least a portion of the sound input corresponds to a predetermined type of sound (e.g., via the sound type detection unit 810), determine whether the sound input includes predetermined content (e.g., via the trigger sound detection unit 812) upon determining that the sound input includes the predetermined content (e.g., via the service initiation unit 814), and initiate a speech-based service (e.g., via the service initiation unit 814) upon determining that the sound input includes the predetermined content.
[0138] In some implementations, processing unit 806 is also configured to determine whether the sound input satisfies a predetermined condition (e.g., with noise detection component 808) before determining whether the sound input corresponds to a predetermined type of sound. In some implementations, processing unit 806 is also configured to determine whether the sound input corresponds to the voice of a particular user (e.g., with voice authentication component 816).
[0139] According to some implementations, FIG. 9 illustrates a functional block diagram of an electronic device 900 configured in accordance with the principles of the present invention as described above. The functional blocks of this device may be implemented in hardware, software, or a combination of hardware and software to implement the principles of the present invention. It will be understood by those skilled in the art that the functional blocks described in FIG. 9 may be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description herein may support any possible combination or division, or definition of additional functional blocks described herein.
[0140] 9, the electronic device 900 includes a voice trigger unit 902. The voice trigger unit 902 can be operated in a variety of different modes. In a first mode, the voice trigger unit receives and determines whether a sound input meets certain criteria (e.g., listening mode). In a second mode, the voice trigger unit 902 does not receive and / or process sound input (e.g., standby mode). The electronic device 900 also includes a processing unit 906 coupled to the voice trigger unit 902. In some implementations, the processing unit 906 includes an environment detector 908 that may include and / or interface with one or more sensors (e.g., including a microphone, camera, accelerometer, gyroscope, etc.) and a mode switching unit 910. In some implementations, the processing unit 906 is configured to determine whether the electronic device is in a substantially enclosed space (e.g., via the environment detection unit 908) by detecting that one or more of the microphone and camera of the electronic device are occluded, and to switch the voice trigger to a second mode (e.g., via the mode switching unit 910) upon determining that the electronic device is in a substantially enclosed space.
[0141] In some implementations, the processing unit is configured to determine whether the electronic device is in a predetermined orientation (e.g., via the environment detection unit 908), and upon determining that the electronic device is in the predetermined orientation, enable a predetermined mode of the voice trigger (e.g., via the mode switching unit 910).
[0142] According to some implementations, FIG. 10 illustrates a functional block diagram of an electronic device 1000 configured in accordance with the principles of the present invention as described above. The functional blocks of this device may be implemented by hardware, software, or a combination of hardware and software to implement the principles of the present invention. It will be understood by those skilled in the art that the functional blocks described in FIG. 10 may be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description herein may support any possible combination or division, or definition of additional functional blocks described herein.
[0143] 10, the electronic device 1000 includes a voice trigger unit 1002. The voice trigger unit 1002 can be operated in a variety of different modes. In a first mode, the voice trigger unit receives sound input and determines whether it meets certain criteria (e.g., listening mode). In a second mode, the voice trigger unit 1002 does not receive and / or process sound input (e.g., standby mode). The electronic device 1000 also includes a processing unit 1006 coupled to the voice trigger unit 1002. In some implementations, the processing unit 1006 includes an environment detector 1008, which may include and / or interface with a microphone and / or camera, and a mode switching unit 1010.
[0144] The processing unit 1006 is configured to determine whether the electronic device is within a substantially enclosed space (e.g., via the environment detection unit 1008) by detecting that one or more of the microphone and camera of the electronic device are occluded, and to switch the voice trigger to a second mode (e.g., via the mode switching unit 1010) upon determining that the electronic device is within a substantially enclosed space. The foregoing description has been set forth with reference to specific embodiments for purposes of illustration. However, the exemplary description above is not intended to be exhaustive or to limit the disclosed embodiments to the precise form. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described in order to best explain the principles and practical applications of the disclosed concepts, thereby enabling those skilled in the art to best utilize the same with various modifications suited to the particular use contemplated.
[0145] It should be understood that, while terms such as "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first sound detector can be referred to as a second sound detector, and similarly, a second sound detector can be referred to as a first sound detector, without changing the meaning of the description, as long as the name is consistently changed for all occurrences of "first sound detector" and "second sound detector" for all occurrences. The first sound detector and the second sound detector are both sound detectors, but are not the same sound detector.
[0146] The terminology used herein is for the purpose of describing particular embodiments and is not intended to limit the scope of the claims. As used in the description of the illustrated embodiments and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. As used herein, it should also be understood that the term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items. It should further be understood that the terms "comprises" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "if" can be interpreted to mean "when" or "upon" or "upon determining" or "in accordance with determining" or "upon detecting" that the aforementioned condition is true, depending on the context. Similarly, the phrases "if it is determined that [the foregoing condition is true]" or "if [the foregoing condition is true]" or "when [the foregoing condition is true]" may be interpreted to mean "upon determining," "upon determining of," or "in response to determining," or "according to determining," or "upon detecting," or "in response to detecting" that the foregoing condition is true.
Claims
1. 1. A method for operating a voice trigger executed on an electronic device including one or more processors and a memory storing instructions for execution by the one or more processors, the method comprising: receiving an audio input; determining whether at least a portion of the sound input corresponds to a predetermined type of sound; determining whether the sound input includes predetermined content upon determining that at least a portion of the sound input corresponds to the predetermined type; initiating a speech-based service upon determining that the sound input includes the predetermined content; A method comprising:
2. 2. The method of claim 1, wherein the step of determining whether the sound input corresponds to a predetermined type of sound is performed by a first sound detector and the step of determining whether the sound input includes predetermined content is performed by a second sound detector, the first sound detector consuming less power during operation than the second sound detector.
3. 3. The method of claim 2, wherein the second sound detector is activated in response to the first sound detector determining that the sound input corresponds to the predetermined type.
4. 3. The method of claim 2, wherein the second sound detector is operated for at least a predetermined time period after the first sound detector determines that the sound input corresponds to the predetermined type.
5. 2. The method of claim 1, wherein the predetermined type is a human voice and the predetermined content is one or more words.
6. 2. The method of claim 1, wherein the predetermined content is one or more predetermined phonemes.
7. 7. The method of claim 6, wherein the one or more predetermined phonemes comprise at least one word.
8. 10. The method of claim 1, further comprising determining whether the sound input satisfies a predetermined condition before determining whether the sound input corresponds to a predetermined type of sound.
9. 9. The method of claim 8, wherein the predetermined condition is an amplitude threshold.
10. 9. The method of claim 8, wherein the step of determining whether the sound input satisfies a predetermined condition is performed by a third sound detector, the third sound detector consuming less power in operation than the first sound detector.
11. storing at least a portion of the sound input in a memory; providing the portion of the sound input to the speech-based service when the speech-based service is initiated; The method of claim 1 further comprising:
12. 10. The method of claim 1, further comprising determining whether the sound input corresponds to the voice of a particular user.
13. 13. The method of claim 12, wherein the speech-based service is initiated upon determining that the sound input includes the predetermined content and that the sound input corresponds to the voice of the particular user.
14. 14. The method of claim 13, wherein upon determining that the sound input includes the predetermined content and that the sound input does not correspond to the voice of the particular user, the speech-based service is initiated in a limited access mode.
15. 14. The method of claim 13, further comprising the step of outputting a voice prompt including the name of the particular user upon determining that the sound input corresponds to the voice of the particular user.
16. determining whether the electronic device is in a predetermined orientation; enabling a predetermined mode of the voice trigger when determining that the electronic device is in the predetermined orientation; The method of claim 1 further comprising:
17. 1. A method for operating a voice trigger executed on an electronic device including one or more processors and a memory storing instructions for execution by the one or more processors, the method comprising: operating a voice trigger in a first mode; determining whether the electronic device is within a substantially enclosed space by detecting that one or more of a microphone and a camera of the electronic device are occluded; switching the voice trigger to a second mode upon determining that the electronic device is within a substantially enclosed space; A method comprising:
18. 18. The method of claim 17, wherein the second mode is a standby mode.
19. 18. The method of claim 17, wherein the first mode is a listening mode.
20. 1. A method for operating a voice trigger executed on an electronic device including one or more processors and a memory storing instructions for execution by the one or more processors, the method comprising: determining whether the electronic device is in a predetermined orientation; enabling a predetermined mode of a voice trigger when determining that the electronic device is in the predetermined orientation; A method comprising:
21. 21. The method of claim 20, wherein the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing downwards, and the predetermined mode is a standby mode.
22. 21. The method of claim 20, wherein the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing upward, and the predetermined mode is a listening mode.
23. 1. A computer-readable storage medium storing one or more programs for execution by one or more processors of an electronic device, the one or more programs comprising: instructions for receiving sound input; instructions for determining whether at least a portion of the sound input corresponds to a predetermined type of sound; instructions for determining whether the sound input includes predetermined content upon determining that at least a portion of the sound input corresponds to the predetermined type; instructions for initiating a speech-based service upon determining that the sound input includes the predetermined content; 1. A computer-readable storage medium comprising:
24. a sound receiving unit configured to receive a sound input; a processing unit coupled to the sound receiving unit; An electronic device comprising: determining whether at least a portion of the sound input corresponds to a predetermined type of sound; determining whether the sound input includes predetermined content upon determining that at least a portion of the sound input corresponds to the predetermined type; Upon determining that the sound input includes the predetermined content, initiate a speech-based service. An electronic device characterized by being configured as follows.
25. 25. The electronic device of claim 24, wherein the processing unit is further configured to determine whether the sound input satisfies a predetermined condition before determining whether the sound input corresponds to a predetermined type of sound.
Citation Information
Patent Citations
Speaker recognizing device
JP2001265385A
Voice recognition device and voice recognition method
JP2012211932A
Device and method for voice recognition in mobile body
JP2012256001A
Interactive speech recognition device and system for hands-free building control
US8340975B1
Multisensory speech detection
WO2010054373A2